In recent years, LLM-based multi-agent systems (MAS) have advanced rapidly, using a router to decompose tasks and delegate subtasks to specialized agents. A natural way to expand capability is to scale up the agent pool by continually integrating new functional agents or tool interfaces, but naive expansion can trigger performance collapse when the router cold-starts on newly added, heterogeneous, and unreliable agents. We propose MonoScale, an expansion-aware update framework that proactively generates a small set of agent-conditioned familiarization tasks, harvests evidence from both successful and failed interactions, and distills it into auditable natural-language memory to guide future routing. We formalize sequential augmentation as a contextual bandit and perform trust-region memory updates, yielding a monotonic non-decreasing performance guarantee across onboarding rounds. Experiments on GAIA and Humanity's Last Exam show stable gains as the agent pool grows, outperforming naive scale-up and strong-router fixed-pool baselines.
翻译:近年来,基于大语言模型的多智能体系统(MAS)发展迅速,通过路由器分解任务并将子任务委托给专门智能体。扩展能力的一种自然方式是不断集成新功能智能体或工具接口以扩大智能体池,但朴素扩展可能导致性能崩溃——当路由器对新添加的、异构且不可靠的智能体进行冷启动时。我们提出MonoScale,一种面向扩展的更新框架,该框架主动生成一小批基于智能体的熟悉化任务,从成功和失败的交互中收集证据,并将其蒸馏为可审计的自然语言记忆以指导未来路由。我们将序贯增强形式化为上下文赌博机,并执行信任区域记忆更新,从而在智能体加入轮次中提供单调非递减的性能保证。在GAIA和"人类最后考试"上的实验表明,随着智能体池扩大,性能稳定提升,优于朴素扩展和强路由器固定池基线方法。