Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets remain largely geographically homogeneous, failing to capture geo-cultural differences. Further, it remains unclear whether such differences persist after controlling for demographics such as age, gender, and ethnicity. Through a meta-analysis of safety datasets, we find that most do not report geo-cultural information, and those that do lack a unified methodology to jointly analyze geo-cultural and demographic correlates. Using the Inglehart-Welzel dimensions of cross-cultural variation, we demonstrate via multilevel modeling that cultural zone membership explains variance in safety ratings beyond standard demographics (p<0.05 across 6 datasets). Moreover, our analysis indicates that roughly 10% of items in the datasets we examined are culturally sensitive: likely to be misclassified as safe without adequate cultural representation. We evaluate LLMs as both rater surrogates and triage tools, finding that current LLMs do not reliably stand in for raters, though they can help prioritize culturally sensitive items for human annotation. Our findings motivate more culturally pluralistic safety evaluation and offer practical takeaways to support it.
翻译:AI模型的安全全球部署需要与因文化而异的人类价值观对齐。然而,安全评估数据集中的评分者池在地理分布上仍高度同质,未能捕捉地理文化差异。此外,在控制年龄、性别和种族等人口统计变量后,这些差异是否依然存在尚不明确。通过对安全数据集的元分析,我们发现大多数数据集未报告地理文化信息,而已报告的数据集也缺乏统一方法论来联合分析地理文化与人口统计相关性。采用英格尔哈特-韦尔策尔跨文化变异维度,我们通过多层模型证明,文化区域归属能解释超越标准人口统计变量的安全评分变异(p<0.05,涵盖6个数据集)。此外,我们的分析表明,所检查数据集中约10%的项目具有文化敏感性:若缺乏充分的文化代表性,这些项目可能被误分类为安全。我们将大语言模型评估为评分者替代工具和分流工具,发现当前大语言模型虽能辅助优先筛选文化敏感项目供人工标注,但无法可靠替代评分者。我们的研究结果推动了更具文化多元性的安全评估实践,并为此提供了实用启示。