Automated short-answer scoring lags other LLM applications. We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies, modeling the traditional effect size of Quadratic Weighted Kappa (QWK) with mixed effects metaregression. We quantitatively illustrate that that the level of difficulty for human experts to perform the task of scoring written work of children has no observed statistical effect on LLM performance. Particularly, we show that some scoring tasks measured as the easiest by human scorers were the hardest for LLMs. Whether by poor implementation by thoughtful researchers or patterns traceable to autoregressive training, on average decoder-only architectures underperform encoders by 0.37--a substantial difference in agreement with humans. Additionally, we measure the contributions of various aspects of LLM technology on successful scoring such as tokenizer vocabulary size, which exhibits diminishing returns--potentially due to undertrained tokens. Findings argue for systems design which better anticipates known statistical shortcomings of autoregressive models. Finally, we provide additional experiments to illustrate wording and tokenization sensitivity and bias elicitation in high-stakes education contexts, where LLMs demonstrate racial discrimination. Code and data for this study are available.
翻译:自动化简答题评分落后于其他大语言模型应用。我们通过对大语言模型简答题评分研究的系统性综述,对890项累积研究结果进行元分析,采用混合效应元回归建模二次加权卡帕系数的传统效应量。我们定量证明:人类专家评分儿童书面作业任务的难度水平,对大语言模型表现并无统计上的显著影响。尤其是,我们发现某些被人类评分者评为最简单的评分任务,对大语言模型而言却最为困难。无论源于研究者的欠佳实施,还是可追溯至自回归训练的模式,解码器架构在人类一致性上平均比编码器低0.37个效应量。此外,我们测量了大语言模型技术中分词器词汇量等要素对成功评分的贡献,其显示出边际效益递减——可能源于未充分训练的token。研究结果主张系统设计应更好预判自回归模型已知的统计缺陷。最后,我们通过附加实验阐明高风险教育情境下的措辞与分词敏感性及偏差诱导现象——大语言模型在此类情境中表现出种族歧视。本研究代码与数据均已公开。