Analysing learning in Multi-Agent Reinforcement Learning (MARL) environments is challenging, in particular with respect to \textit{individual} decision-making. Practitioners frequently struggle to compare training runs due to the inherent stochasticity in algorithms arising from random dithering exploration, environment transition noise, and stochastic gradient updates to name a few. Traditional analytical approaches, such as replicator dynamics, oft rely on mean-field approximations to remove stochastic effects, but this simplification, whilst able to provide general overall trends, can lead to dissonance between analytical predictions and actual agent realisations. We propose modelling MARL training as a \textit{coupled stochastic dynamical systems}, capturing both agent interactions and environmental characteristics. Leveraging tools from dynamical systems theory, we pragmatically analyse the stability and sensitivity of agent behaviour, which are key dimensions for their practical deployments, for example, in presence of strict safety requirements. This framework allows us to rigorously study the inherent stochasticity of MARL, providing a deeper understanding of system behaviour.


翻译:多智能体强化学习环境中的学习分析具有挑战性,特别是关于个体决策的层面。由于算法固有的随机性(包括随机抖动探索、环境转移噪声及随机梯度更新等),实践者常难以比较不同训练过程。传统分析方法(如复制动力学)通常依赖平均场近似来消除随机效应,但这种简化虽能揭示整体趋势,却可能导致分析预测与实际智能体实现之间存在偏差。我们提出将多智能体强化学习训练建模为耦合随机动力系统,同时捕捉智能体交互与环境特征。借助动力系统理论工具,我们务实分析了智能体行为的稳定性与敏感性——这是其实际部署(例如存在严格安全要求时)的关键维度。该框架使我们能够严谨研究多智能体强化学习内在的随机性,从而更深入地理解系统行为。

0
下载
关闭预览

相关内容

《多智能体强化学习中的机制设计优化研究》103页
专知会员服务
34+阅读 · 2025年5月31日
多智能体强化学习控制与决策研究综述
专知会员服务
50+阅读 · 2024年11月23日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
多智能体深度强化学习研究进展
专知会员服务
76+阅读 · 2024年7月17日
多智能体博弈学习研究进展
专知会员服务
91+阅读 · 2024年5月5日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
多智能体协同决策方法研究
专知会员服务
136+阅读 · 2022年12月15日
「基于通信的多智能体强化学习」 进展综述
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
《多智能体强化学习中的机制设计优化研究》103页
专知会员服务
34+阅读 · 2025年5月31日
多智能体强化学习控制与决策研究综述
专知会员服务
50+阅读 · 2024年11月23日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
多智能体深度强化学习研究进展
专知会员服务
76+阅读 · 2024年7月17日
多智能体博弈学习研究进展
专知会员服务
91+阅读 · 2024年5月5日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
多智能体协同决策方法研究
专知会员服务
136+阅读 · 2022年12月15日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员