Journal ArticleParallel publicationPublished versionDOI: 10.48548/pubdata-4151

A comprehensive item pretesting study to develop an automated short-answer grading platform for Turkish inputs

Chronological data

Date of first publication2026-08-21
Date of publication in PubData 2026-08-24

Language of the resource

English

Related external resources

Variant form of DOI: 10.1007/s44217-026-02033-4
Kılıç, M., Cüvitoğlu, G., Polan, Ş., Balcı, Y., Gürel, S., Demir, E., Akşehirli, S., Namdar, B., Karakaş, N., Metin, S., Atılgan, H., Kışla, T., & Aydın, B. (2026). A comprehensive item pretesting study to develop an automated short-answer grading platform for Turkish inputs. Discover Education, 5(1), Article 755.
Published in ISSN: 2731-5525
Discover Education

Abstract

This study presents several pretesting approaches for automated short-answer grading (ASAG) items, addressing a critical gap in standardized item development procedures in the AS literature. We utilized a large language model approach to develop and assess an ASAG platform for Turkish language inputs across three subject domains: history of the Turkish revolution, Turkish as a first language, and science literacy. Our methodology integrated traditional classical test theory item parameters with insights from cognitive interviews and expert-based machine score-ability ratings to predict ASAG performance. We evaluated performance using quadratic weighted kappa across 73 short-answer items. Our preliminary results indicated that the effectiveness of the pretesting metrics varies by domain. While strong correlations between item parameters and ASAG performance were observed in the history and Turkish domains, a more complex case arose in the science literacy domain, where pretest metrics were insignificant predictors, likely due to linguistic and semantic diversity and complex reasoning. Preliminary evidence suggests that, depending on the study domain, ASAG performance might be predicted from pretesting results. Overall, this study has the potential to advance the field by shifting the paradigm from traditional, retrospective algorithmic validation to a proactive, evidence-based process of item pretesting and development for ASAG.

Keywords

Automated Short-answer Grading (ASAG); Item Pretesting; Validity; Machine Score-ability; Natural Language Processing

Leuphana Institution

More information

DDC

Creation Context

Research