On-the-fly Directed Controller Synthesis (OTF-DCS) mitigates state-space explosion by incrementally exploring the system and relies critically on an exploration policy to guide search efficiently. Recent reinforcement learning (RL) approaches learn such policies and achieve promising zero-shot generalization from small training instances to larger unseen ones. However, a fundamental limitation is anisotropic generalization, where an RL policy exhibits strong performance only in a specific region of the domain-parameter space while remaining fragile elsewhere due to training stochasticity and trajectory-dependent bias. To address this, we propose a Soft Mixture-of-Experts framework that combines multiple RL experts via a prior-confidence gating mechanism and treats these anisotropic behaviors as complementary specializations. The evaluation on the Air Traffic benchmark shows that Soft-MoE substantially expands the solvable parameter space and improves robustness compared to any single expert.


翻译:在线定向控制器综合通过增量式探索系统来缓解状态空间爆炸问题,其关键依赖于探索策略以高效引导搜索。最近的强化学习方法通过学习此类策略,在从小型训练实例到未见大型实例的零样本泛化方面展现出良好前景。然而,一个根本性局限在于各向异性泛化:由于训练随机性和轨迹依赖性偏差,强化学习策略仅在域参数空间的特定区域表现优异,而在其他区域则表现脆弱。为解决此问题,我们提出一种软专家混合框架,该框架通过先验置信度门控机制融合多个强化学习专家,并将这些各向异性行为视为互补的专业化能力。在空管基准测试上的评估表明,相较于任何单一专家,软专家混合框架显著扩展了可求解参数空间并提升了鲁棒性。

0
下载
关闭预览

相关内容

【CMU博士论文】基于课程学习的鲁棒强化学习
专知会员服务
21+阅读 · 2025年3月27日
【CMU博士论文】通过课程学习实现鲁棒的强化学习
专知会员服务
25+阅读 · 2024年12月15日
面向强化学习的可解释性研究综述
专知会员服务
45+阅读 · 2024年7月30日
《图强化学习在组合优化中的应用》综述
专知会员服务
61+阅读 · 2024年4月10日
【加州理工博士论文】基于学习的鲁棒控制方法,137页pdf
专知会员服务
32+阅读 · 2023年12月23日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证
深度强化学习实验室
20+阅读 · 2020年8月11日
基于逆强化学习的示教学习方法综述
计算机研究与发展
16+阅读 · 2019年2月25日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关VIP内容
【CMU博士论文】基于课程学习的鲁棒强化学习
专知会员服务
21+阅读 · 2025年3月27日
【CMU博士论文】通过课程学习实现鲁棒的强化学习
专知会员服务
25+阅读 · 2024年12月15日
面向强化学习的可解释性研究综述
专知会员服务
45+阅读 · 2024年7月30日
《图强化学习在组合优化中的应用》综述
专知会员服务
61+阅读 · 2024年4月10日
【加州理工博士论文】基于学习的鲁棒控制方法,137页pdf
专知会员服务
32+阅读 · 2023年12月23日
相关基金
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员