As multi-agent systems (MAS) become increasingly complex, identifying the contributions of individual agents is critical for system optimization. However, existing approaches lack a rigorous, unified framework for credit assignment. In this work, we formalize agent attribution as a cooperative game, parameterized by the coalition distribution, removal protocol, and target metric. Using this framework, we show that Leave-One-Out (LOO) identifies bottleneck agents as effectively as combinatorial methods, but at a fraction of the computational cost. We also demonstrate that removal protocols induce distinct games: Agent ablation isolates structural bottlenecks, whereas introspective LLM judges fail to faithfully approximate this behavior. Furthermore, to evaluate the utility of specific agent backbones, we introduce attribution via model replacement. By substituting underlying models of low-contribution agents, we improve task performance by up to 17% while reducing cost by up to 35% across three benchmarks. Finally, we apply our framework to audit a medical MAS, revealing that agent contributions to diagnostic accuracy and ethical behavior are often decoupled. By intervening on counterproductive roles, we observe an increase in ethics alignment while maintaining diagnostic accuracy. Overall, this work provides a principled approach for cost-effective MAS attribution and intervention.
翻译:随着多智能体系统(MAS)日益复杂,识别单个智能体的贡献对系统优化至关重要。然而,现有方法缺乏严谨统一的信用分配框架。本文提出将智能体归因形式化为合作博弈模型,该模型由联盟分布、消融协议和目标指标参数化。基于该框架,我们证明留一法(LOO)能以组合方法相同效果识别瓶颈智能体,但计算成本仅为其极小部分。此外,我们揭示消融协议会诱发不同博弈行为:智能体消融可识别结构性瓶颈,而内省式大模型评判器无法准确逼近该行为。为评估特定智能体骨干网络效用,我们引入模型替换归因方法。通过替换低贡献智能体的底层模型,我们在三个基准测试中实现任务性能提升最高17%,同时成本降低最高35%。最后,将该框架应用于医疗多智能体系统审计,发现智能体对诊断准确性与伦理行为的贡献往往相互解耦。通过干预反生产角色,我们观察到在维持诊断准确性的同时提升了伦理对齐度。总体而言,本研究为经济高效的多智能体系统归因与干预提供了原则性方法。