Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- until recently there has been limited work on the design and development of datasets for cultural assessment. Here, we review existing approaches to such datasets and identify their main limitations. To address these issues, we propose design guidelines for annotators and report on the construction of a dataset built according to these principles. We further present a series of contrastive experiments conducted with this dataset. The results demonstrate that our design yields test sets with greater discriminative power, effectively distinguishing between models specialized for a given culture and those that are not, ceteris paribus.
翻译:尽管大型语言模型的文化(不)适配性问题——常被表述为文化偏见——已引起广泛关注,但迄今为止,针对文化评估数据集设计与开发的相关研究仍较为有限。本文系统梳理了现有文化评估数据集的研究方法,并识别其主要局限性。针对这些问题,我们提出了面向标注者的设计准则,并报告了依据这些原则构建的数据集建设过程。进一步地,我们利用该数据集开展了一系列对比实验。结果表明,我们的设计方案能够生成更具区分度的测试集,在保持其他条件不变的情况下,有效识别出针对特定文化进行专门优化的模型与未优化模型之间的差异。