Conventional optimization-based metering depends on strict adherence to precomputed schedules, which limits the flexibility required for the stochastic operations of Advanced Air Mobility (AAM). In contrast, multi-agent reinforcement learning (MARL) offers a decentralized, adaptive framework that can better handle uncertainty, required for safe aircraft separation assurance. Despite this advantage, current MARL approaches often overfit to specific airspace structures, limiting their adaptability to new configurations. To improve generalization, we recast the MARL problem in a relative polar state space and train a transformer encoder model across diverse traffic patterns and intersection angles. The learned model provides speed advisories to resolve conflicts while maintaining aircraft near their desired cruising speeds. In our experiments, we evaluated encoder depths of 1, 2, and 3 layers in both structured and unstructured airspaces, and found that a single encoder configuration outperformed deeper variants, yielding near-zero near mid-air collision rates and shorter loss-of-separation infringements than the deeper configurations. Additionally, we showed that the same configuration outperforms a baseline model designed purely with attention. Together, our results suggest that the newly formulated state representation, novel design of neural network architecture, and proposed training strategy provide an adaptable and scalable decentralized solution for aircraft separation assurance in both structured and unstructured airspaces.


翻译:传统的基于优化的流量管理依赖于对预先计算的时间表的严格遵循,这限制了先进空中交通(AAM)随机运行所需的灵活性。相比之下,多智能体强化学习(MARL)提供了一种去中心化、自适应的框架,能更好地处理不确定性,而这正是安全飞机间隔保障所必需的。尽管有此优势,当前的MARL方法常常过度拟合特定的空域结构,限制了其对新配置的适应性。为提升泛化能力,我们在相对极坐标状态空间中重新构建了MARL问题,并针对多种交通模式和交叉角度训练了一个Transformer编码器模型。学习到的模型提供速度建议以解决冲突,同时使飞机保持在接近其期望巡航速度的状态。在我们的实验中,我们在结构化和非结构化空域中评估了1、2和3层的编码器深度,发现单一编码器配置优于更深层的变体,相较于更深层的配置,产生了接近零的近空中碰撞率和更短的间隔丧失违规时间。此外,我们证明了相同配置优于纯粹基于注意力设计的基线模型。总之,我们的结果表明,新构建的状态表示、新颖的神经网络架构设计以及提出的训练策略,为结构化和非结构化空域中的飞机间隔保障提供了一种适应性强且可扩展的去中心化解决方案。

0
下载
关闭预览

相关内容

多智能体强化学习中的稳健且高效的通信
专知会员服务
26+阅读 · 2025年11月17日
开放环境下的协作多智能体强化学习进展综述
专知会员服务
35+阅读 · 2025年1月19日
《空战战术多智能体强化学习中的可解释性》最新报告
专知会员服务
86+阅读 · 2024年10月25日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
「博弈论视角下多智能体强化学习」研究综述
专知会员服务
185+阅读 · 2022年4月30日
「基于通信的多智能体强化学习」 进展综述
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
16+阅读 · 2020年9月9日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
综述 | Self-Evolving Coding Agents:自进化编程智能体
专知会员服务
0+阅读 · 今天13:16
美海军陆战队将三型无人机整合入统一战场网络
专知会员服务
2+阅读 · 今天9:39
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 今天9:12
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 今天9:08
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 今天8:46
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
2+阅读 · 今天8:40
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
相关VIP内容
多智能体强化学习中的稳健且高效的通信
专知会员服务
26+阅读 · 2025年11月17日
开放环境下的协作多智能体强化学习进展综述
专知会员服务
35+阅读 · 2025年1月19日
《空战战术多智能体强化学习中的可解释性》最新报告
专知会员服务
86+阅读 · 2024年10月25日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
「博弈论视角下多智能体强化学习」研究综述
专知会员服务
185+阅读 · 2022年4月30日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员