Analysing learning in Multi-Agent Reinforcement Learning (MARL) environments is challenging, in particular with respect to \textit{individual} decision-making. Practitioners frequently struggle to compare training runs due to the inherent stochasticity in algorithms arising from random dithering exploration, environment transition noise, and stochastic gradient updates to name a few. Traditional analytical approaches, such as replicator dynamics, oft rely on mean-field approximations to remove stochastic effects, but this simplification, whilst able to provide general overall trends, can lead to dissonance between analytical predictions and actual agent realisations. We propose modelling MARL training as a \textit{coupled stochastic dynamical systems}, capturing both agent interactions and environmental characteristics. Leveraging tools from dynamical systems theory, we pragmatically analyse the stability and sensitivity of agent behaviour, which are key dimensions for their practical deployments, for example, in presence of strict safety requirements. This framework allows us to rigorously study the inherent stochasticity of MARL, providing a deeper understanding of system behaviour.
翻译:多智能体强化学习环境中的学习分析具有挑战性,特别是关于个体决策的层面。由于算法固有的随机性(包括随机抖动探索、环境转移噪声及随机梯度更新等),实践者常难以比较不同训练过程。传统分析方法(如复制动力学)通常依赖平均场近似来消除随机效应,但这种简化虽能揭示整体趋势,却可能导致分析预测与实际智能体实现之间存在偏差。我们提出将多智能体强化学习训练建模为耦合随机动力系统,同时捕捉智能体交互与环境特征。借助动力系统理论工具,我们务实分析了智能体行为的稳定性与敏感性——这是其实际部署(例如存在严格安全要求时)的关键维度。该框架使我们能够严谨研究多智能体强化学习内在的随机性,从而更深入地理解系统行为。