Testing standards for AI-based scores in automated essay scoring.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42536698.
- Also identified by DOI 10.1371/journal.pone.0354680.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Recent developments in the field of artificial intelligence and machine learning allow the wide application of large language models for the evaluation of written text and other non-numerical data. When applied in the context of psychological and educational assessments, such models can be used for assigning scores to essays and other types of responses. In contrast to classical tests, essays do not consist of test items, which leads to specific challenges in the evaluation of testing standards for scores obtained from AI models that differ from those observed for classical ability tests and personality questionnaires. To address these challenges, we discuss the evaluation of validity, fairness, and reliability for scores obtained from models of artificial intelligence in the context of automated essay scoring. We review existing methods, propose new methods, and further illustrate the reviewed methods with an empirical example based on the Hewlett Foundation data set on automated essay scoring. By applying the proposed framework to an evaluation based on a DistilBERT model, we find the model to be robust with sufficiently high internal consistency (Spearman-Brown coefficients in the range from .77 to .92). We further found empirical evidence for the validity of the evaluation model, but also indications for violations of fairness when comparing the human and AI scores across different topics. This study provides a standardized, replicable toolkit for researchers and practitioners to evaluate the psychometric quality of AI-based assessments.
Medical subject headings
- Artificial Intelligence
- Educational Measurement