Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents. Decentralised shields can enforce safety at runtime, but purely factorised permissions often exclude optimal team behaviour that is safe only through coordination. We study deterministic safety guarantees for agents trained and deployed under decentralised execution, recovering team-optimal safe behaviour without centralised runtime control. Agents have a shared global specification $φ$ in the safety fragment of Linear Temporal Logic ($\mathsf{LTL}_{\mathsf{safe}}$ ), and select among tuples of local $\mathsf{LTL}_{\mathsf{safe}}$ obligations whose conjunction implies the global specification $φ$. Each agent may rely on the other agents' local obligations as assumptions because the whole contract tuple is certified simultaneously and allows projection into local action masks. At learning time, a non-stationary multi-armed bandit chooses among a library of local $\mathsf{LTL}_{\mathsf{safe}}$ obligations to select the tuple that optimises team reward, all without forgoing end-to-end safety. We evaluate the approach across 6 environments and 15 algorithmic variants.


翻译:在全局安全无法由任何智能体单方面强制实施时,多智能体强化学习中出现安全协调问题:一个智能体动作的可行性可能依赖于其他智能体的动态过程。去中心化屏蔽能在运行时强制执行安全性,但纯粹的因子化权限通常会排除通过协调才能实现安全的、最优的团队行为。我们研究了在去中心化执行下训练和部署的智能体的确定性安全保证,在无需中心化运行时控制的情况下恢复了团队最优的安全行为。智能体共享共享一个以线性时序逻辑($\mathsf{LTL}_{\mathsf{safe}}$)安全片段表述的全局规范$φ$,并选择一组局部$\mathsf{LTL}_{\mathsf{safe}}$义务的元组,这些义务的合取蕴含全局规范$φ$。每个智能体可将其他智能体的局部义务作为假设依赖,因为整个合约元组是同时认证的,并允许投影到局部动作掩码中。在训练时,非平稳多臂赌博机从本地$\mathsf{LTL}_{\mathsf{safe}}$义务库中选择元组以优化团队奖励,且全程不放弃端到端安全性。我们在6个环境和15种算法变体上评估了该方法。

0
下载
关闭预览

相关内容

智能体,顾名思义,就是具有智能的实体,英文名是Agent。
多智能体强化学习中的稳健且高效的通信
专知会员服务
26+阅读 · 2025年11月17日
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
42+阅读 · 2025年4月18日
开放环境下的协作多智能体强化学习进展综述
专知会员服务
35+阅读 · 2025年1月19日
多智能体强化学习控制与决策研究综述
专知会员服务
50+阅读 · 2024年11月23日
基于多智能体强化学习的博弈综述
专知会员服务
53+阅读 · 2024年11月23日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
多智能体协同决策方法研究
专知会员服务
136+阅读 · 2022年12月15日
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
智能合约的形式化验证方法研究综述
专知
16+阅读 · 2021年5月8日
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
多智能体强化学习中的稳健且高效的通信
专知会员服务
26+阅读 · 2025年11月17日
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
42+阅读 · 2025年4月18日
开放环境下的协作多智能体强化学习进展综述
专知会员服务
35+阅读 · 2025年1月19日
多智能体强化学习控制与决策研究综述
专知会员服务
50+阅读 · 2024年11月23日
基于多智能体强化学习的博弈综述
专知会员服务
53+阅读 · 2024年11月23日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
多智能体协同决策方法研究
专知会员服务
136+阅读 · 2022年12月15日
相关资讯
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
智能合约的形式化验证方法研究综述
专知
16+阅读 · 2021年5月8日
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员