This paper compares historical annotations by humans and Large Language Models. The findings reveal that both exhibit some cultural bias, but Large Language Models achieve a higher consensus on the interpretation of historical facts from short texts. While humans tend to disagree on the basis of their personal biases, Large Models disagree when they skip information or produce hallucinations. These findings have significant implications for digital humanities, enabling large-scale annotation and quantitative analysis of historical data. This offers new educational and research opportunities to explore historical interpretations from different Language Models, fostering critical thinking about bias.
翻译:本文比较了人类与大语言模型对历史文本的标注结果。研究发现,两者均表现出一定的文化偏见,但大语言模型在对短文本历史事实的解释上展现出更高的一致性。人类倾向于基于个人偏见产生分歧,而大语言模型则因信息遗漏或幻觉生成而产生分歧。这一发现对数字人文学科具有重要意义,可实现历史数据的大规模标注与量化分析。通过探索不同语言模型的历史解读方式,本研究为教育及研究提供了新机遇,有助于培养对偏见的批判性思考。