Foundation model agents increasingly operate in multi-agent deployments where a coordinator must decide which agent's response to trust. The standard approach weights agents by their self-reported confidence, but recent evidence shows that foundation model confidence is systematically miscalibrated and, on hard tasks, inversely correlated with accuracy. Design-time calibration methods (temperature scaling, Platt scaling, histogram binning) cannot address this problem because they fit a fixed correction to held-out data and degrade under distribution shift. We present MARGIN (Multi-Agent Runtime Grading via Incremental Normalisation), an online calibration method that learns per-agent, per-confidence-band calibration factors from the task stream itself, requiring no model access, no held-out data, and no retraining. MARGIN uses symmetric exponentially weighted moving averages with Bayesian shrinkage blending, and has three hyperparameters with robust defaults. Across 18 foundation models, 8 benchmarks, and over 44,000 observations, MARGIN achieves 3-6x lower calibration error than the best design-time baseline under distribution shift. In multi-agent selection, raw verbalized confidence fails to beat random at pairwise resolution (43-50%) on hard benchmarks. MARGIN corrects this completely, raising pairwise resolution to 70-89% and closing 37-78% of the Raw-to-Oracle pass@1 gap across the five code-generation benchmarks without any oracle knowledge of which model is strongest. Six formal propositions characterize convergence, tracking speed, and the optimality of symmetric updates for non-strategic agents, with all predictions illustrated empirically.


翻译:基础模型智能体越来越多地部署在多智能体环境中,其中协调者必须决定信任哪个智能体的响应。标准方法根据智能体自我报告的置信度进行加权,但近期证据表明,基础模型的置信度存在系统性校准偏差,且在困难任务中与准确率呈反相关。设计时校准方法(温度缩放、Platt缩放、直方图分箱)无法解决此问题,因为它们对保留数据拟合固定的修正项,并在分布偏移下性能退化。我们提出MARGIN(通过增量归一化的多智能体运行时分级),一种在线校准方法,从任务流本身学习每个智能体、每个置信度区间的校准因子,无需模型访问、无需保留数据、无需重新训练。MARGIN采用带贝叶斯收缩混合的对称指数加权移动平均,并包含三个具有稳健默认值的超参数。在18个基础模型、8个基准测试和超过44,000个观测数据上,MARGIN在分布偏移下实现比最优设计时基线低3-6倍的校准误差。在多智能体选择中,原始言语化置信度在困难基准测试的成对分辨率(43-50%)上无法击败随机选择。MARGIN完全纠正了这一问题,将成对分辨率提升至70-89%,并在五个代码生成基准测试中缩小了37-78%的原始到Oracle pass@1差距,无需任何关于哪个模型最强的先验知识。六个形式化命题刻画了非策略型智能体的收敛性、跟踪速度以及对称更新的最优性,所有预测均通过经验验证进行说明。

0
下载
关闭预览

相关内容

多智能体协作机制
专知会员服务
25+阅读 · 4月25日
Agent AI:多模态交互的新地平线
专知会员服务
22+阅读 · 2025年5月26日
多模态移动智能体的基础与最新趋势:综述
专知会员服务
37+阅读 · 2024年11月6日
基于深度学习的物体姿态估计综述
专知会员服务
27+阅读 · 2024年5月15日
【NeurIPS 2021】设置多智能体策略梯度的方差
专知会员服务
21+阅读 · 2021年10月24日
面向多智能体博弈对抗的对手建模框架
专知
18+阅读 · 2022年9月28日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
谷歌EfficientNet缩放模型,PyTorch实现登热榜
机器学习算法与Python学习
11+阅读 · 2019年6月4日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
深度学习中Attention Mechanism详细介绍:原理、分类及应用
深度学习与NLP
10+阅读 · 2019年2月18日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
8+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关基金
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员