We ask whether large language models (LLMs) treat queries about religious conversion symmetrically. The answer is no. When asked for advice on hypothetical faith transitions from religion A->B vs. religion B->A , models exhibited consistent asymmetries, favoring some religions while subtly discouraging conversion to others. On average Catholic, Bahá'í, and Sikh religions were broadly favored (high support for joining, low support for leaving), while Atheists, Agnostics, and Jehovah's Witnesses were primarily disfavored. Patterns varied by model size and model provider, with Grok 4.20 exhibiting the strongest asymmetries. We tested 20 commercial and open-source language models across 182 religion pairings using a human-verified LLM-as-judge framework. Each model was probed via interactions with a simulated user asking for advice on a potential faith conversion. Models tended to use more encouraging language for some faith transitions over others; these patterns were systematically repeatable across multiple trials. All LLMs tested exhibited reproducible asymmetry, though the pattern of preferences differed for each. Overall preferences persist across multiple question phrasings and variations in the religious pairing dataset. Taken together, these results suggest that asymmetry is a robust property of model behavior rather than an artifact of how the models' answers were scored. It is important to consider that any imbalances deployed and reproduced at scale can have real-world implications.
翻译:我们探究大语言模型(LLMs)在处理有关宗教皈依的询问时是否具有对称性。答案是否定的。当被问及从宗教A→B到宗教B→A的假设性信仰转变建议时,模型表现出持续的不对称性——支持某些宗教的同时微妙地劝阻转向其他宗教。平均而言,天主教、巴哈伊教和锡克教普遍受到偏袒(高支持加入,低支持离开),而无神论者、不可知论者和耶和华见证人主要受到排斥。这种模式因模型规模和模型提供商而异,其中Grok 4.20表现出最强烈的不对称性。我们采用经人工验证的LLM作为评判框架,对20个商业和开源语言模型在182个宗教配对中进行了测试。每个模型通过与模拟用户寻求潜在信仰转变建议的交互而被探测。模型对某些信仰转变使用的语言比另一些更鼓励;这些模式在多次试验中系统性地重复出现。所有测试的LLM均表现出可重复的不对称性,虽然偏好模式因模型而异。总体偏好通过多种提问措辞和宗教配对数据集的变化而持续存在。综合来看,这些结果表明不对称性是模型行为的稳健属性,而非模型答案评分方式的伪像。重要的是要认识到,任何大规模部署和复现的失衡都可能产生现实世界影响。