Recent work has shown that LLMs overrepresent dominant cultures, particularly Western ones, while marginalizing others. We investigate whether this affects models' ability to generate culturally adapted responses by evaluating their use of local measurement units based on the user's perceived cultural background. We introduce Cultural and Pragmatic Response Inference (CAPRI), a dataset of conversations with varying levels of cultural cues. Experiments with state-of-the-art LLMs show that models can infer cultural background and recall relevant conventions, but often fail to utilize the information to adapt their answers to the relevant cultural conventions, unless explicitly prompted to perform the tasks sequentially. We further evaluate adaptation to the interpretation of time and quantity expressions, two subjective language grounding dimensions that are affected by culture. We find that models increasingly adapt their answers as cultural cues accumulate, but their priors are not culture-neutral, sometimes aligning with the model's country of origin. Overall, CAPRI provides a resource for future research aimed at narrowing the gap between cultural knowledge and culturally adaptive language generation.
翻译:近期研究表明,大型语言模型过度代表主流文化(尤其是西方文化),同时边缘化其他文化。我们通过评估模型基于用户感知文化背景使用本地计量单位的能力,探究这一现象是否影响模型生成文化适应性回应的能力。我们提出了文化与语用回应推断(CAPRI)数据集,其中包含具有不同文化线索级别的对话。对最先进的大型语言模型的实验表明,模型能够推断文化背景并回忆相关惯例,但往往未能利用这些信息将其回应调整至相应文化惯例,除非明确提示其按顺序执行任务。我们进一步评估了模型对时间和数量表达解释的适应能力——这是两个受文化影响的主观语言基础维度。研究发现,随着文化线索的累积,模型会逐渐调整其回应,但其先验知识并非文化中立,有时会与模型来源国的文化倾向一致。总体而言,CAPRI为未来旨在缩小文化知识与文化适应性语言生成之间差距的研究提供了资源。