LLMs are deployed globally, yet produce responses biased towards cultures with abundant training data. Existing cultural localization approaches such as prompting or post-training alignment are black-box, hard to control, and do not reveal whether failures reflect missing knowledge or poor elicitation. In this paper, we address these gaps using mechanistic interpretability to uncover and manipulate cultural representations in LLMs. Leveraging sparse autoencoders, we identify interpretable features that encode culturally salient information and aggregate them into Cultural Embeddings (CuE). We use CuE both to analyze implicit cultural biases under underspecified prompts and to construct white-box steering interventions. Across multiple models, we show that CuE-based steering increases cultural faithfulness and elicits significantly rarer, long-tail cultural concepts than prompting alone. Notably, CuE-based steering is complementary to black-box localization methods, offering gains when applied on top of prompt-augmented inputs. This also suggests that models do benefit from better elicitation strategies, and don't necessarily lack long-tail knowledge representation, though this varies across cultures. Our results provide both diagnostic insight into cultural representations in LLMs and a controllable method to steer towards desired cultures.
翻译:大型语言模型(LLMs)虽已实现全球部署,但其输出结果仍偏向训练数据丰富的文化群体。现有文化本地化方法(如提示工程或训练后对齐)存在黑箱特性、难以操控,且无法区分模型失败的原因是缺乏知识表征还是知识提取能力不足。本文通过机械可解释性方法揭示并操控LLMs中的文化表征以弥补上述不足。我们利用稀疏自编码器识别编码文化显著性信息的可解释特征,并将其聚合为文化嵌入向量(Cultural Embeddings, CuE)。CuE既可分析模糊提示下的隐性文化偏见,也可构建白盒调控干预机制。跨模型实验表明,相较于单纯使用提示工程,基于CuE的调控方法能显著提升文化忠实度,并有效生成更罕见的长尾文化概念。值得注意的是,CuE调控与黑箱本地化方法具有互补性——在提示增强输入基础上叠加CuE调控可进一步提升性能。这亦表明模型确实能通过优化知识提取策略获益,且未必缺乏长尾知识表征(尽管不同文化群体存在差异)。本研究既为LLMs中的文化表征提供了诊断性洞见,也提供了可控的文化偏好调控方法。