Iterative generative modeling techniques, such as flow matching, provide powerful tools to model complex behaviors for effective offline reinforcement learning (RL). In this work, we propose a new off-policy RL algorithm that trains a flow policy based on prior data. Our idea starts from the "expanded" Markov decision process (MDP) framework, which treats individual flow refinement steps as separate actions in an MDP. To enable off-policy RL within this framework, we apply two techniques: we generate virtual on-policy trajectories (by "reversing" flows) to make this framework compatible with prior data, and we apply a bias-and-variance reduction technique to mitigate the curse of horizon in off-policy RL. We call the resulting algorithm Reversal Q-learning (RQL). RQL has several advantages over previous flow-based RL methods: it does not suffer from backpropagation through time, makes better use of the learned value function, and directly trains the full, expressive flow policy. Through our experiments on 50 challenging simulated robotic tasks, we show that RQL leads to the best average offline RL performance compared to state-of-the-art flow-based offline RL algorithms.


翻译:迭代式生成建模技术(如流匹配)为高效离线强化学习(RL)中的复杂行为建模提供了强大工具。本文提出一种基于先验数据训练流策略的新型离策略强化学习算法。我们的核心思路源于“扩展”马尔可夫决策过程(MDP)框架,该框架将流细化步骤视为MDP中的独立动作。为在此框架中实现离策略强化学习,我们应用了两种技术:通过“逆向”流生成虚拟在策略轨迹,使框架与先验数据兼容;同时采用偏差-方差缩减技术缓解离策略强化学习中的视界诅咒。我们将由此产生的算法命名为逆向Q学习(Reversal Q-learning, RQL)。与以往基于流的强化学习方法相比,RQL具有多项优势:无需沿时间反向传播、能更充分利用学习到的价值函数、可直接训练完整且表达能力强的流策略。通过在50个具有挑战性的模拟机器人任务上的实验表明,与最先进的基于流的离线强化学习算法相比,RQL实现了最优的平均离线强化学习性能。

0
下载
关闭预览

相关内容

逆向强化学习研究综述*
专知会员服务
60+阅读 · 2023年10月13日
基于模型的强化学习综述
专知会员服务
48+阅读 · 2023年1月9日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
强化学习开篇:Q-Learning原理详解
AINLP
37+阅读 · 2020年7月28日
元强化学习迎来一盆冷水:不比元Q学习好多少
AI科技评论
12+阅读 · 2020年2月27日
入门 | 通过 Q-learning 深入理解强化学习
机器之心
12+阅读 · 2018年4月17日
【强化学习】强化学习/增强学习/再励学习介绍
产业智能官
10+阅读 · 2018年2月23日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Arxiv
0+阅读 · 5月10日
VIP会员
最新内容
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
0+阅读 · 37分钟前
《战略战术化:一项综合性述评》
专知会员服务
0+阅读 · 41分钟前
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 今天8:46
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
0+阅读 · 今天8:40
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员