What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL? - 专知论文

会员服务 ·

0

泛化理论 · 可辨认的 · 相互独立的 · state-of-the-art · ENJOY ·

2023 年 5 月 30 日

What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?

翻译：离线目标条件强化学习中实现未见目标泛化的关键要素

Rui Yang,Yong Lin,Xiaoteng Ma,Hao Hu,Chongjie Zhang,Tong Zhang

from arxiv, Accepted by Proceedings of the 40th International Conference on Machine Learning, 2023

Offline goal-conditioned RL (GCRL) offers a way to train general-purpose agents from fully offline datasets. In addition to being conservative within the dataset, the generalization ability to achieve unseen goals is another fundamental challenge for offline GCRL. However, to the best of our knowledge, this problem has not been well studied yet. In this paper, we study out-of-distribution (OOD) generalization of offline GCRL both theoretically and empirically to identify factors that are important. In a number of experiments, we observe that weighted imitation learning enjoys better generalization than pessimism-based offline RL method. Based on this insight, we derive a theory for OOD generalization, which characterizes several important design choices. We then propose a new offline GCRL method, Generalizable Offline goAl-condiTioned RL (GOAT), by combining the findings from our theoretical and empirical studies. On a new benchmark containing 9 independent identically distributed (IID) tasks and 17 OOD tasks, GOAT outperforms current state-of-the-art methods by a large margin.

翻译：离线目标条件强化学习（GCRL）提供了一种从完全离线数据集中训练通用智能体的方法。除了在数据集中保持保守性外，实现未见目标的泛化能力是离线GCRL面临的另一个基本挑战。然而，据我们所知，该问题尚未得到充分研究。本文从理论和实验两方面研究了离线GCRL的分布外（OOD）泛化问题，旨在识别重要影响因素。在多项实验中，我们观察到加权模仿学习比基于悲观主义的离线强化学习方法具有更好的泛化性。基于这一发现，我们推导出OOD泛化理论，该理论刻画了若干关键设计选择。进而，结合理论与实验研究成果，我们提出了一种新的离线GCRL方法——泛化离线目标条件强化学习（GOAT）。在包含9个独立同分布（IID）任务和17个OOD任务的新基准测试中，GOAT的性能大幅领先当前最先进方法。

0

相关内容

泛化理论

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【硬核课】机器人学习课程，UT Austin朱玉可博士讲述自主机器人的人工智能与机器学习机器学习算法

【硬核课】机器人学习课程，UT Austin朱玉可博士讲述自主机器人的人工智能与机器学习机器学习算法

专知会员服务

40+阅读 · 2020年9月21日

因果图，Causal Graphs，52页ppt

因果图，Causal Graphs，52页ppt

专知会员服务

254+阅读 · 2020年4月19日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

84+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

SIRT7的稳定性与DNA损伤修复的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于“痰浊瘀血”理论从TLR/NF-κB信号通路探讨当归芍药散抗阿尔茨海默病(AD)的作用机制

国家自然科学基金

0+阅读 · 2013年12月31日

海洋环境下多无人艇协同航行机制与仿人自适应控制研究

国家自然科学基金

0+阅读 · 2013年12月31日

切换系统的有限时间稳定与镇定研究

国家自然科学基金

0+阅读 · 2012年12月31日

控制力矩陀螺的高频微振动特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

CRMP-2调控弥漫性轴索损伤后海马Aβ表达及学习记忆功能的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

严重低血糖新生大鼠脑内代谢物的高分辨魔角旋转磁共振波谱研究

国家自然科学基金

0+阅读 · 2011年12月31日

乳腺癌化疗所致记忆障碍的脑机制及其康复的研究

国家自然科学基金

0+阅读 · 2011年12月31日

Wnt/β-catenin通路在内皮抑素抑制动脉粥样硬化斑块内微血管新生中的作用及其机制

国家自然科学基金

0+阅读 · 2011年12月31日

基于等价输入干扰的扰动抑制鲁棒控制方法

国家自然科学基金

0+阅读 · 2009年12月31日

Impatient Bandits: Optimizing for the Long-Term Without Delay

Arxiv

0+阅读 · 2023年7月19日

The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions

Arxiv

0+阅读 · 2023年7月19日

Unsupervised Video Anomaly Detection with Diffusion Models Conditioned on Compact Motion Representations

Arxiv

0+阅读 · 2023年7月19日

How Curvature Enhance the Adaptation Power of Framelet GCNs

Arxiv

0+阅读 · 2023年7月19日

Online Learning with Costly Features in Non-stationary Environments

Arxiv

0+阅读 · 2023年7月18日

The Predicted-Deletion Dynamic Model: Taking Advantage of ML Predictions, for Free

Arxiv

0+阅读 · 2023年7月17日

Dynamic neighbourhood optimisation for task allocation using multi-agent

Arxiv

102+阅读 · 2022年5月11日

Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks

Arxiv

10+阅读 · 2022年2月10日

Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

Arxiv

11+阅读 · 2021年9月3日

Attribute-Guided Adversarial Training for Robustness to Natural Perturbations

Arxiv

15+阅读 · 2020年12月3日

VIP会员

文章信息

相关主题

相互独立的

state-of-the-art

最新内容

美国从乌克兰无人机战争中学习经验

美国从乌克兰无人机战争中学习经验

专知会员服务

6+阅读 · 6月21日

ICML 2026 | 面向视觉语言模型的语义鲁棒性认证

ICML 2026 | 面向视觉语言模型的语义鲁棒性认证

专知会员服务

2+阅读 · 6月21日

综述 | 智能体电子设计自动化：从“交接有效性”重新理解Agentic EDA

综述 | 智能体电子设计自动化：从“交接有效性”重新理解Agentic EDA

专知会员服务

4+阅读 · 6月21日

深入解读 Palantir AIP：全球最具争议的人工智能平台究竟如何运作

深入解读 Palantir AIP：全球最具争议的人工智能平台究竟如何运作

专知会员服务

19+阅读 · 6月20日

ICML 2026 | 多任务贝叶斯上下文学习：让 Transformer 在测试时显式适应新先验

ICML 2026 | 多任务贝叶斯上下文学习：让 Transformer 在测试时显式适应新先验

专知会员服务

5+阅读 · 6月19日

ACL 2026综述 | 大规模手语数据集：资源、基准与标注标准

ACL 2026综述 | 大规模手语数据集：资源、基准与标注标准

专知会员服务

8+阅读 · 6月19日

ICML 2026 Spotlight | SmoothSMoE：解析稀疏 MoE 路由不连续

ICML 2026 Spotlight | SmoothSMoE：解析稀疏 MoE 路由不连续

专知会员服务

7+阅读 · 6月18日

综述 | 周期表视角下的大模型推理：范式、方法与失败模式

综述 | 周期表视角下的大模型推理：范式、方法与失败模式

专知会员服务

9+阅读 · 6月18日

《廉价自杀式无人机战争的军事战略影响：乌克兰和伊朗案例研究》

《廉价自杀式无人机战争的军事战略影响：乌克兰和伊朗案例研究》

专知会员服务

13+阅读 · 6月18日

《面向反无人机作战的联邦式可解释射频–光电/红外情报融合：边缘人工智能优化、电子战韧性及分布式监视验证》

《面向反无人机作战的联邦式可解释射频–光电/红外情报融合：边缘人工智能优化、电子战韧性及分布式监视验证》

专知会员服务

12+阅读 · 6月18日

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

ICML 2026 | FR3D：解耦自车运动的未来动态三维重建世界模型

专知会员服务

8+阅读 · 6月17日

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

【伯克利博士论文】迈向可扩展与自我演进的大语言模型智能体

专知会员服务

13+阅读 · 6月17日

学习数据的几何：形状空间分析数学综述

学习数据的几何：形状空间分析数学综述

专知会员服务

10+阅读 · 6月17日

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

《现代防空系统综述：架构、传感器、拦截器及新兴威胁环境对基础设施受限防御环境的影响》2026最新长综述

专知会员服务

24+阅读 · 6月17日

定向能反无人机系统最新发展动态

定向能反无人机系统最新发展动态

专知会员服务

12+阅读 · 6月17日

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【硬核课】机器人学习课程，UT Austin朱玉可博士讲述自主机器人的人工智能与机器学习机器学习算法

【硬核课】机器人学习课程，UT Austin朱玉可博士讲述自主机器人的人工智能与机器学习机器学习算法

专知会员服务

40+阅读 · 2020年9月21日

因果图，Causal Graphs，52页ppt

因果图，Causal Graphs，52页ppt

专知会员服务

254+阅读 · 2020年4月19日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

84+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

ICML 2026 | 面向视觉语言模型的语义鲁棒性认证

深入解读 Palantir AIP：全球最具争议的人工智能平台究竟如何运作

美国从乌克兰无人机战争中学习经验

综述 | 智能体电子设计自动化：从“交接有效性”重新理解Agentic EDA

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Impatient Bandits: Optimizing for the Long-Term Without Delay

Arxiv

0+阅读 · 2023年7月19日

The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions

Arxiv

0+阅读 · 2023年7月19日

Unsupervised Video Anomaly Detection with Diffusion Models Conditioned on Compact Motion Representations

Arxiv

0+阅读 · 2023年7月19日

How Curvature Enhance the Adaptation Power of Framelet GCNs

Arxiv

0+阅读 · 2023年7月19日

Online Learning with Costly Features in Non-stationary Environments

Arxiv

0+阅读 · 2023年7月18日

The Predicted-Deletion Dynamic Model: Taking Advantage of ML Predictions, for Free

Arxiv

0+阅读 · 2023年7月17日

Dynamic neighbourhood optimisation for task allocation using multi-agent

Arxiv

102+阅读 · 2022年5月11日

Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks

Arxiv

10+阅读 · 2022年2月10日

Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

Arxiv

11+阅读 · 2021年9月3日

Attribute-Guided Adversarial Training for Robustness to Natural Perturbations

Arxiv

15+阅读 · 2020年12月3日

相关基金

SIRT7的稳定性与DNA损伤修复的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于“痰浊瘀血”理论从TLR/NF-κB信号通路探讨当归芍药散抗阿尔茨海默病(AD)的作用机制

国家自然科学基金

0+阅读 · 2013年12月31日

海洋环境下多无人艇协同航行机制与仿人自适应控制研究

国家自然科学基金

0+阅读 · 2013年12月31日

切换系统的有限时间稳定与镇定研究

国家自然科学基金

0+阅读 · 2012年12月31日

控制力矩陀螺的高频微振动特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

CRMP-2调控弥漫性轴索损伤后海马Aβ表达及学习记忆功能的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

严重低血糖新生大鼠脑内代谢物的高分辨魔角旋转磁共振波谱研究

国家自然科学基金

0+阅读 · 2011年12月31日

乳腺癌化疗所致记忆障碍的脑机制及其康复的研究

国家自然科学基金

0+阅读 · 2011年12月31日

Wnt/β-catenin通路在内皮抑素抑制动脉粥样硬化斑块内微血管新生中的作用及其机制

国家自然科学基金

0+阅读 · 2011年12月31日

基于等价输入干扰的扰动抑制鲁棒控制方法

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员