Automated Essay Scoring automates the grading process of essays, providing a great advantage for improving the writing proficiency of students. While holistic essay scoring research is prevalent, a noticeable gap exists in scoring essays for specific quality traits. In this work, we focus on the relevance trait, which measures the ability of the student to stay on-topic throughout the entire essay. We propose a novel approach for graded relevance scoring of written essays that employs dense retrieval encoders. Dense representations of essays at different relevance levels then form clusters in the embeddings space, such that their centroids are potentially separate enough to effectively represent their relevance levels. We hence use the simple 1-Nearest-Neighbor classification over those centroids to determine the relevance level of an unseen essay. As an effective unsupervised dense encoder, we leverage Contriever, which is pre-trained with contrastive learning and demonstrated comparable performance to supervised dense retrieval models. We tested our approach on both task-specific (i.e., training and testing on same task) and cross-task (i.e., testing on unseen task) scenarios using the widely used ASAP++ dataset. Our method establishes a new state-of-the-art performance in the task-specific scenario, while its extension for the cross-task scenario exhibited a performance that is on par with the state-of-the-art model for that scenario. We also analyzed the performance of our approach in a more practical few-shot scenario, showing that it can significantly reduce the labeling cost while sacrificing only 10% of its effectiveness.
翻译:自动作文评分将作文评分过程自动化,为提升学生写作能力提供了巨大优势。尽管整体性作文评分研究已十分普遍,但在针对特定质量特征进行评分方面仍存在显著空白。本文聚焦于相关性这一特征——它衡量学生在整篇作文中保持主题一致的能力。我们提出了一种基于稠密检索编码器的书面作文分级相关性评分新方法。不同相关性等级的作文的稠密表示在嵌入空间中形成聚类,其聚类中心可能充分分离,从而有效表征各等级的相关性。因此,我们采用简单的1-最近邻分类方法对聚类中心进行分类,以判定未见过作文的相关性等级。作为一种有效的无监督稠密编码器,我们利用了Contriever——该模型通过对比学习进行预训练,并展现了与监督式稠密检索模型相当的性能。我们在广泛使用的ASAP++数据集上,分别在任务特定场景(即在同一任务上训练和测试)和跨任务场景(即在未见任务上测试)下测试了所提方法。在任务特定场景中,我们的方法达到了新的最优性能;而在跨任务场景中,其扩展方法的表现与该场景下的最优模型持平。此外,我们分析了所提方法在更实用的少样本场景中的表现,结果显示该方法能够在牺牲仅10%有效性的前提下显著降低标注成本。