Multi-agent systems often require agents to collaborate with or compete against other agents with diverse goals, behaviors, or strategies. Agent modeling is essential when designing adaptive policies for intelligent machine agents in multiagent systems, as this is the means by which the ego agent understands other agents' behavior and extracts their meaningful policy representations. These representations can be used to enhance the ego agent's adaptive policy which is trained by reinforcement learning. However, existing agent modeling approaches typically assume the availability of local observations from other agents (modeled agents) during training or a long observation trajectory for policy adaption. To remove these constrictive assumptions and improve agent modeling performance, we devised a Contrastive Learning-based Agent Modeling (CLAM) method that relies only on the local observations from the ego agent during training and execution. With these observations, CLAM is capable of generating consistent high-quality policy representations in real-time right from the beginning of each episode. We evaluated the efficacy of our approach in both cooperative and competitive multi-agent environments. Our experiments demonstrate that our approach achieves state-of-the-art on both cooperative and competitive tasks, highlighting the potential of contrastive learning-based agent modeling for enhancing reinforcement learning.
翻译:多智能体系统通常要求智能体与具有不同目标、行为或策略的其他智能体进行协作或竞争。在面向多智能体系统的自适应策略设计中,智能体建模至关重要——这是自利智能体理解其他智能体行为并提取其有意义的策略表征的方法。这些表征可用于增强经强化学习训练的自利智能体自适应策略。然而,现有智能体建模方法通常假设在训练期间可获取其他智能体(被建模智能体)的局部观测,或需要较长的观测轨迹进行策略适应。为消除这些限制性假设并提升智能体建模性能,我们设计了基于对比学习的智能体建模方法(CLAM),该方法在训练和执行阶段仅依赖自利智能体的局部观测。借助这些观测数据,CLAM能够从每个回合初始阶段实时生成一致的高质量策略表征。我们在协作与竞争型多智能体环境中验证了该方法的有效性。实验结果表明,我们的方法在协作与竞争任务中均达到最优性能,凸显了基于对比学习的智能体建模对强化学习性能增强的潜力。