A comprehensive item pretesting study to develop an automated short-answer grading platform for Turkish inputs
Chronological data
Date of first publication2026-08-21
Date of publication in PubData 2026-08-24
Language of the resource
English
Editor
Case provider
Other contributors
Abstract
This study presents several pretesting approaches for automated short-answer grading (ASAG) items, addressing a critical gap in standardized item development procedures in the AS literature. We utilized a large language model approach to develop and assess an ASAG platform for Turkish language inputs across three subject domains: history of the Turkish revolution, Turkish as a first language, and science literacy. Our methodology integrated traditional classical test theory item parameters with insights from cognitive interviews and expert-based machine score-ability ratings to predict ASAG performance. We evaluated performance using quadratic weighted kappa across 73 short-answer items. Our preliminary results indicated that the effectiveness of the pretesting metrics varies by domain. While strong correlations between item parameters and ASAG performance were observed in the history and Turkish domains, a more complex case arose in the science literacy domain, where pretest metrics were insignificant predictors, likely due to linguistic and semantic diversity and complex reasoning. Preliminary evidence suggests that, depending on the study domain, ASAG performance might be predicted from pretesting results. Overall, this study has the potential to advance the field by shifting the paradigm from traditional, retrospective algorithmic validation to a proactive, evidence-based process of item pretesting and development for ASAG.
Keywords
Automated Short-answer Grading (ASAG); Item Pretesting; Validity; Machine Score-ability; Natural Language Processing
