Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods either require per-country preference data and fine-tuning budgets or assume white-box access to model internals that commercial APIs do not expose. In this work, we focus on this realistic black-box, public-data-only regime and observe that within-country sociodemographic disagreement, not consensus, is the primary steering signal. We introduce DISCA (Disagreement-Informed Steering for Cultural Alignment), an inference-time method that instantiates each country as a panel of World-Values-Survey-grounded persona agents and converts their disagreement into a bounded, loss-averse logit correction. Across 20 countries and 7 open-weight backbones (2B--70B), DISCA reduces cultural misalignment on MultiTP by 10--24% on the six backbones >=3.8B, and 2--7% on open-ended scenarios, without changing any weights. Our results suggest that inference-time calibration is a scalable alternative to fine-tuning for serving the long tail of global moral preferences.
翻译:大型语言模型日益介入依赖道德判断的决策过程,但大量证据表明其隐含偏好并非文化中立。现有文化对齐方法要么需要各国偏好数据与微调预算,要么假设具备商业API未开放的模型白盒访问权限。本研究聚焦于现实的黑盒、仅使用公开数据场景,发现国内社会人口分歧(而非共识)才是主要的引导信号。我们提出DISCA(分歧引导的文化对齐校准),一种推理时方法:将每个国家实例化为基于世界价值观调查的人格代理组,将其分歧转化为有界且损失规避的logit校正。基于20个国家、7个开放权重基础模型(2B-70B)的实验表明,在≥3.8B参数的六个基础模型上,DISCA将MultiTP上的文化错位减少10-24%,在开放式场景中减少2-7%,且不改变任何权重。我们的结果表明,推理时校正是服务于全球道德偏好长尾分布的微调可扩展替代方案。