As LLM-based multi-agent systems (MAS) are deployed in the wild, the resilience of their collaboration structures against adversarial compromise becomes a critical safety concern. Attackers may leverage prompt-injection or jailbreaking to sabotage individual agents within MAS workflows, but the interaction between model scaling and system-level resilience remains poorly understood. This paper investigates how model scale affects the security of linear multi-agent workflows. Our experiments across scales of two open-weight model families on the HumanEval benchmark reveal a compliance-correction symmetry: larger models are far more likely to faithfully execute malicious instructions, with the control-to-malicious performance drop reaching 53.7pp at 27B in uncorrected pipelines. However, appending a lightweight terminal Fixer stage collapses this to 0.6pp and restores statistical parity with control-level performance, demonstrating that strictly linear collaboration structures can be viable and resilient to adversaries at this scale, and suggesting that the brittleness previously attributed to linear topology may stem from a lack of correction.
翻译:随着基于大语言模型的多智能体系统(MAS)在实际环境中部署,其协作结构在面临对抗性妥协时的鲁棒性成为关键的安全问题。攻击者可能通过提示注入或越狱手段破坏MAS工作流中的单个智能体,但模型规模扩展与系统级鲁棒性之间的相互作用仍鲜为人知。本文研究模型规模如何影响线性多智能体工作流的安全性。我们在HumanEval基准测试上,针对两个开源模型家族的不同规模进行实验,揭示了一种“顺从-修正对称性”:规模更大的模型更容易忠实执行恶意指令,在未修正流程中,控制组与恶意组性能差异最大达53.7个百分点(27B参数规模)。然而,附加一个轻量级终端修复阶段后,该差异缩小至0.6个百分点,并恢复与控制组性能的统计等价性,这表明严格的线性协作结构在该规模下具有生存力与抗对抗性,且此前归因于线性拓扑结构的脆弱性可能源于缺乏修正机制。