Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic probability axioms even when every component is locally coherent. We formalise this locally coherent, globally incoherent failure via the compositional residual eps*, the L2 distance from the composed quote to the joint coherent polytope, computable at runtime from system output and the declared cross-component coupling constraints. A product-structure dichotomy characterises when local coherence suffices, and a Rayleigh-quotient prediction matches the observed residual within 7% on three of four relation classes. A hierarchical Boyle-Dykstra projection repairs the composition deterministically; an anytime-valid e-process gives sequential coherence monitoring. Across 1,876 ensemble cliques on a four-LLM mid-tier panel (frontier-panel rerun in Section 5.5), eps* > 0 on 33-94% of cliques, translating to +0.115 nats per bet of regret on 1,770 resolved bets under the proportional allocation rule (the gain collapses to +0.006 under bettors that themselves coherentise). Three intuitive LLM-side mitigations(retrieval, partition-aware prompting, aggregator-LLM) each fail or regress.


翻译:多组件大语言模型智能体将各组件提供的概率性主张进行组合,而每个组件仅能观测联合问题的局部信息;即便每个组件在局部范围内满足相干性,组合结果仍可能违反基本概率公理。我们通过组合残差ε*(组合命题与联合相干多面体之间的L2距离,可在运行时根据系统输出及声明的跨组件耦合约束计算)形式化描述了这种"局部相干、整体不相干"的失效模式。乘积结构二分法刻画了局部相干性足以保证全局一致性的条件,而瑞利商预测方法在四类关系中的三类上实现了与观测残差7%以内的匹配精度。分层Boyle-Dykstra投影算法能以确定性方式修复组合结果;任意有效的e过程实现序贯相干性监控。在包含四组大语言模型的中端模型面板(前沿模型面板重测见第5.5节)产生的1,876个集成团中,33%-94%的团存在ε*>0的情况,在比例分配规则下对应1,770个已结算赌注中每注+0.115纳特的遗憾值(当使用自身已相干化的投注者时,增益降至+0.006)。三种直观的大语言模型端缓解策略(检索增强、分区感知提示、聚合型大语言模型)均告失败或出现性能倒退。

0
下载
关闭预览

相关内容

多智能体协作机制
专知会员服务
27+阅读 · 4月25日
《多智能体大语言模型系统的可靠决策研究》
专知会员服务
43+阅读 · 2月2日
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
大语言模型与小语言模型协同机制综述
专知会员服务
40+阅读 · 2025年5月15日
多循环嵌套的大语言模型多智能体指挥控制过程
专知会员服务
45+阅读 · 2025年1月19日
大模型报告:模型能力决定下限,场景适配度决定上限
专知会员服务
57+阅读 · 2024年6月3日
面向多智能体博弈对抗的对手建模框架
专知
19+阅读 · 2022年9月28日
数据受限条件下的多模态处理技术综述
专知
22+阅读 · 2022年7月16日
【深度语义匹配模型】原理篇二:交互篇
AINLP
16+阅读 · 2020年5月18日
专家报告|深度学习+图像多模态融合
中国图象图形学报
12+阅读 · 2019年10月23日
自然语言处理(NLP)知识结构总结
AI100
51+阅读 · 2018年8月17日
概率图模型体系:HMM、MEMM、CRF
机器学习研究会
30+阅读 · 2018年2月10日
各种相似性度量及Python实现
机器学习算法与Python学习
11+阅读 · 2017年7月6日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
Arxiv
0+阅读 · 5月28日
VIP会员
最新内容
反制无人机:乌克兰提供的五点启示
专知会员服务
7+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
7+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
4+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
8+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
10+阅读 · 9月22日
战争不仅需要机器人:人类仍不可或缺
专知会员服务
6+阅读 · 9月21日
《描绘美国防部创新基础设施的未来蓝图》100页
专知会员服务
10+阅读 · 9月21日
相关资讯
面向多智能体博弈对抗的对手建模框架
专知
19+阅读 · 2022年9月28日
数据受限条件下的多模态处理技术综述
专知
22+阅读 · 2022年7月16日
【深度语义匹配模型】原理篇二:交互篇
AINLP
16+阅读 · 2020年5月18日
专家报告|深度学习+图像多模态融合
中国图象图形学报
12+阅读 · 2019年10月23日
自然语言处理(NLP)知识结构总结
AI100
51+阅读 · 2018年8月17日
概率图模型体系:HMM、MEMM、CRF
机器学习研究会
30+阅读 · 2018年2月10日
各种相似性度量及Python实现
机器学习算法与Python学习
11+阅读 · 2017年7月6日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员