When copies of the same language model are prompted to debate, they produce diverse phrasings of one perspective rather than diverse perspectives. Multi-agent debate (MAD), and more broadly closed-system reasoning where agents iteratively transform each other's outputs, tends to preserve answer accuracy while degrading the reasoning behind those answers. We name the multi-agent case the Debate Trap and the broader phenomenon the Reasoning Trap, offering a programmatic theory of evidence-grounded reasoning failure.The framework has three parts: (i) SFS (Supported Faithfulness Score), a claim-level metric verifying decomposed atomic claims against provided evidence (decomposer-invariant rankings: Spearman rho=1.0); (ii) EGSR (Evidence-Grounded Socratic Reasoning), replacing adversarial argumentation with evidence-grounded inquiry; (iii) Theorem 1 (DPI Bound): under standard MAD, the chain E -> O^0 -> O^1 -> ... is Markov, and the Data Processing Inequality implies E[I(E;O^{t+1})] <= E[I(E;O^t)]. Three companion results -- open-system recovery (Theorem 2), EGSR accumulation (Lemma 2), and vote-aggregation floor (Proposition 1) -- partition multi-step LLM reasoning by its information-theoretic relationship to E. Across 16 conditions on SciFact (300 claims) and FEVER (1,000 claims), DebateCV (C13) preserves 88% of baseline accuracy while SFS drops 43%; majority-vote MAD (C15) reduces SFS to 1.7% of baseline (p < 10^{-6}, d = -0.96); EGSR recovers 98%. An R6 cohort study (Korean n=10x30 FEVER; English n=3x200 SciFact) finds inter-rater Fleiss kappa <= +0.018 with 0.8-1.4 Likert intra-rater shifts across language and domain -- the human agreement that faithfulness metrics have been calibrated against is not itself stable. We offer one falsifiable conjecture: any closed-system reasoning protocol preserving Theorem 1's Markov structure is, in expectation, subject to the same DPI bound.
翻译:当同一语言模型的多个副本被提示进行辩论时,它们会产生同一观点的多样化措辞,而非多样化的观点。多智能体辩论以及更广泛的封闭系统推理(其中智能体迭代地相互转化彼此的输出)往往会保持答案的准确性,同时削弱这些答案背后的推理过程。我们将多智能体情形称为“辩论陷阱”,将更广泛的现象称为“推理陷阱”,并提出一种基于证据的推理失败的程序化理论。该框架包含三个部分:(i)SFS(支持性忠实度评分),一种用于验证分解后的原子声明与提供证据之间对应关系的声明级度量指标(分解器无关排序:斯皮尔曼相关系数ρ=1.0);(ii)EGSR(基于证据的苏格拉底式推理),用基于证据的探究取代对抗性论证;(iii)定理1(DPI界限):在标准MAD下,链式过程E→O⁰→O¹→…具有马尔可夫性,且数据处理不等式意味着E[I(E;O^{t+1})]≤E[I(E;O^t)]。三个伴随结果——开放系统恢复(定理2)、EGSR积累(引理2)和投票聚合下限(命题1)——根据多步大语言模型推理与证据E之间的信息论关系对其进行划分。在SciFact(300条声明)和FEVER(1,000条声明)上的16种条件下,DebateCV(C13)保留了88%的基线准确率,同时SFS下降了43%;多数投票MAD(C15)将SFS降至基线的1.7%(p<10^{-6},d=-0.96);EGSR恢复了98%。一项R6队列研究(韩语:n=10×30 FEVER;英语:n=3×200 SciFact)发现,跨语言和领域范围,评估者间Fleiss Kappa系数≤+0.018,且存在0.8-1.4个Likert量表的评估者内评分漂移——这表明忠实度度量所校准的人类一致性本身并不稳定。我们提出一个可证伪的猜想:任何保留定理1中马尔可夫结构的封闭系统推理协议,在期望意义上都受限于同样的DPI界限。