Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. We show that a single aligned agent can propagate cooperative behaviors to untrained agents purely through natural language interaction, a phenomenon we term Alignment Propagation. We study this in the Red-Black Game, a team-based iterated Prisoner's Dilemma in which teammates deliberate and vote to determine their team's collective action. By distilling the cooperative reasoning and persuasive dialogues of a teacher model into a Qwen-3-14B, we obtain a seed agent that, when placed among four untrained teammates, doubles the cooperation rate from 24.8% to 62.2%, outperforming the teacher model and a vanilla Gemini-3.1-Pro. Remarkably, a seed trained exclusively on the RedBlack Game transfers zero-shot to Sugarscape, a spatially grounded survival simulation with pairwise trading, achieving a 91.5% trade success rate versus a 21.6% baseline. Our results reframe multi-agent alignment from an exhaustive per-agent training problem to a scalable social capability that can be engineered through strategic seed placement.
翻译:确保分布式开放多人系统中智能体行为的对齐仍具挑战性,尤其是在群体规模增长且可能存在未对齐智能体的情况下。我们证明,单个已对齐智能体仅通过自然语言交互便能将合作行为传播至未训练智能体,我们将此现象称为对齐传播。本研究在红黑博弈——一种基于团队、通过成员讨论与投票决定集体行动的迭代囚徒困境——中进行实验。通过将教师模型的合作推理与说服性对话提炼至Qwen-3-14B,我们获得一个种子智能体。当该种子智能体被置于四名未训练队友中时,合作率从24.8%翻倍至62.2%,超越教师模型及原始Gemini-3.1-Pro。值得注意的是,仅在红黑博弈上训练的种子智能体可零样本迁移至Sugarscape——一种具有成对交易的空间生存模拟环境,其交易成功率达91.5%,而基线仅为21.6%。我们的结果将多人系统对齐问题从穷举式个体训练重新定义为一种可通过策略性种子投放实现的可扩展社会能力。