Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation. We ask whether the polarity of prior conversation history biases subsequent judgments, an effect we call the accumulated message effect on LLM judgments (AMEL). Across 84,088 API calls to 12 models from 5 providers (OpenAI, Anthropic, Google, DeepSeek, and four open-source models), we present identical test items in isolation or following histories saturated with predominantly positive or negative evaluations. Models shift toward the conversation's prevailing polarity (d = -0.17, p < 10^-53). The effect concentrates on items where the model is genuinely uncertain at baseline (d = -0.36 for high-entropy items, vs d = -0.15 when the baseline is deterministic). Bias does not grow with context length: 5 prior turns and 50 produce the same shift (Spearman |r| < 0.01; OLS slope p = 0.80). And there is a negativity asymmetry: paired per item, negative histories induce 1.52x more bias than positive (t = 13.03, p < 10^-36, n = 2,733). Scaling helps but does not solve it (Anthropic: Haiku -0.22 to Opus -0.17; OpenAI: Nano -0.34 to GPT-5.2 -0.17). Three follow-ups narrow the mechanism. The token probability distribution shifts continuously, not at a threshold. The negativity asymmetry has both token-level and semantic components, though attributing the balance is exploratory at our sample sizes. Position does not matter: five biased turns anywhere in a 50-turn history produce the same shift. The simplest fix for evaluation pipelines is a fresh context per item; when batching is unavoidable, balancing the history helps.


翻译:大语言模型被常规用作自动评估者:用于审查代码、审核内容或评分输出,且往往通过同一对话处理大量条目。我们探究对话历史的情感极性是否会偏倚后续判断——我们将此效应称为LLM判断的累积消息影响(AMEL)。通过对来自5个提供商(OpenAI、Anthropic、Google、DeepSeek及四个开源模型)的12个模型进行84,088次API调用,我们在隔离条件下或紧随以正面或负面评估为主导的历史对话后,呈现相同的测试项目。模型会向对话的主流情感极性偏移(d=-0.17,p<10^-53)。该效应集中于那些在基线状态下模型确实不确定的项目(高熵项目d=-0.36,而确定性基线项目d=-0.15)。偏倚不随上下文长度增长:5轮先验对话与50轮产生相同的偏移(Spearman |r|<0.01;OLS斜率p=0.80)。此外存在负性不对称:按项目配对,负面历史诱导的偏倚是正面历史的1.52倍(t=13.03,p<10^-36,n=2,733)。模型规模扩大有助于缓解但无法解决该问题(Anthropic:Haiku -0.22至Opus -0.17;OpenAI:Nano -0.34至GPT-5.2 -0.17)。三项后续实验缩小了机制范围。标记概率分布是连续偏移而非阈值式变化。负性不对称同时包含标记级与语义成分,但在我们的样本量下平衡归因仅为探索性分析。位置无关紧要:在50轮历史中的任意位置插入五轮偏倚对话均产生相同偏移。评估流水线最简单的修复方案是为每个项目使用全新上下文;当批处理不可避免时,平衡历史对话有所助益。

0
下载
关闭预览

相关内容

智能体评判者(Agent-as-a-Judge)研究综述
专知会员服务
37+阅读 · 1月9日
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
Llama-3-SynE:实现有效且高效的大语言模型持续预训练
专知会员服务
36+阅读 · 2024年7月30日
自然语言处理顶会EMNLP2018接受论文列表!
专知
87+阅读 · 2018年8月26日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
《最强大的军事网状网络》
专知会员服务
3+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
3+阅读 · 9月7日
分层反无人机系统发展新趋势
专知会员服务
10+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员