RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process. We present Durable Evaluation Framework (DEF) Arbitration, a multi-agent architecture that mitigates identity-framed sycophancy by arbitrating between two models tuned to opposing DEFs, with a pragmatist synthesizer evaluating both arguments blind to their origins. This paper evaluates a prompt-based instantiation of DEF Arbitration. The key mechanisms are static DEF tuning, identity stripping before synthesis, single-round independent argumentation, and blind arbitration. We evaluate five instantiations on 200 stratified questions from SycophancyEval. All tested DEF variants (AnCifer, DeWin, FeynStein, BurGal, Trident) significantly outperform the single-model baseline (18.5%) and instructed-opposition baseline (29.0%), with DeWin achieving 48.5% accuracy (z=6.36, p<0.001 versus both). The variants are not significantly different from each other at n=200. The BurGal variant achieves 53.0% but functions as an architectural validity check; its consensus/heterodox axis structurally favors the heterodox model on every benchmark question. A pre-training floor affects an estimated 40% of questions; fine-tuned DEF models are the identified next step.
翻译:RLHF训练的模型在准确性上存在系统性偏向于一致性的结构特性。我们提出持久评估框架(DEF)仲裁架构——一种多智能体系统,通过协调两个针对对立DEF调优的模型来缓解身份框架性谄媚,并由实用主义综合器在不考虑来源的情况下对双方论点进行评估。本文对该架构的提示式实现进行了评估。关键技术机制包括:静态DEF调优、综合前的身份剥离、单轮独立论证以及盲式仲裁。我们在SycophancyEval的200个分层问题上测试了五种实现方式。所有DEF变体(AnCifer、DeWin、FeynStein、BurGal、Trident)均显著优于单模型基线(18.5%)和指令对立基线(29.0%),其中DeWin达到48.5%的准确率(与两者比较:z=6.36, p<0.001)。在n=200条件下各变体间无显著差异。BurGal变体虽实现53.0%的准确率,但主要作为架构有效性校验存在——其共识/异见轴在每个基准问题上都结构性偏向异见模型。预训练底层限制影响约40%的问题,微调型DEF模型被确定为下一步研究方向。