Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in Healthcare - 专知论文

会员服务 ·

0

分解的 · Better · Learning · 推断 · 强化学习 ·

2023 年 5 月 2 日

Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in Healthcare

翻译：利用因子化动作空间实现医疗领域高效离线强化学习

Shengpu Tang,Maggie Makar,Michael W. Sjoding,Finale Doshi-Velez,Jenna Wiens

from arxiv, 30 pages, 18 figures, 2 tables. NeurIPS 2022. Code available at https://github.com/MLD3/OfflineRL_FactoredActions

Many reinforcement learning (RL) applications have combinatorial action spaces, where each action is a composition of sub-actions. A standard RL approach ignores this inherent factorization structure, resulting in a potential failure to make meaningful inferences about rarely observed sub-action combinations; this is particularly problematic for offline settings, where data may be limited. In this work, we propose a form of linear Q-function decomposition induced by factored action spaces. We study the theoretical properties of our approach, identifying scenarios where it is guaranteed to lead to zero bias when used to approximate the Q-function. Outside the regimes with theoretical guarantees, we show that our approach can still be useful because it leads to better sample efficiency without necessarily sacrificing policy optimality, allowing us to achieve a better bias-variance trade-off. Across several offline RL problems using simulators and real-world datasets motivated by healthcare, we demonstrate that incorporating factored action spaces into value-based RL can result in better-performing policies. Our approach can help an agent make more accurate inferences within underexplored regions of the state-action space when applying RL to observational datasets.

翻译：许多强化学习应用涉及组合型动作空间，其中每个动作由子动作组合而成。标准强化学习方法忽略了这种固有的分解结构，导致对罕见子动作组合的推理可能失效——这在数据可能有限的离线场景中尤其成问题。本文提出一种由因子化动作空间诱导的线性Q函数分解方法。我们研究了该方法的理论性质，识别了在使用其近似Q函数时保证零偏差的场景。在理论保证范围之外，我们证明该方法仍具价值：它在不必然牺牲策略最优性的前提下提升样本效率，从而实现更优的偏差-方差权衡。通过利用医疗领域激励的仿真器与真实世界数据集进行多个离线强化学习实验，我们证明将因子化动作空间融入基于价值的强化学习可生成更优策略。该方法能帮助智能体在将强化学习应用于观测数据集时，对状态-动作空间未被充分探索的区域做出更准确的推理。

0

相关内容

分解的

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

专知会员服务

247+阅读 · 2019年10月21日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

37+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

61+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

60+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

164+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

疏肝法对情绪调节不良MCI患者工作记忆影响的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

离子取代与晶格调控对YAG:Ce3+荧光粉发光性质的影响

国家自然科学基金

0+阅读 · 2014年12月31日

平流层突然增温期间电离层赤道异常区周期性变化研究

国家自然科学基金

0+阅读 · 2014年12月31日

FRP加固钢筋混凝土柱受压性能的尺寸效应研究及工程应用

国家自然科学基金

0+阅读 · 2013年12月31日

慢性应激对肿瘤术后复发和转移的影响及机制探讨

国家自然科学基金

0+阅读 · 2013年12月31日

MDSCs在动脉粥样硬化中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

Pharicin B稳定维甲酸受体的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

随机变分不等式

国家自然科学基金

0+阅读 · 2011年12月31日

基于HM耦合效应的高渗透水压水工隧洞承载机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

离子类溶质在土中迁移过程的耦合效应仿真分析

国家自然科学基金

0+阅读 · 2009年12月31日

Robotic Packaging Optimization with Reinforcement Learning

Arxiv

0+阅读 · 2023年6月16日

Semi-Offline Reinforcement Learning for Optimized Text Generation

Arxiv

0+阅读 · 2023年6月16日

Generalizable Resource Scaling of 5G Slices using Constrained Reinforcement Learning

Arxiv

0+阅读 · 2023年6月15日

Neuroevolution is a Competitive Alternative to Reinforcement Learning for Skill Discovery

Arxiv

0+阅读 · 2023年6月15日

RLTP: Reinforcement Learning to Pace for Delayed Impression Modeling in Preloaded Ads

Arxiv

0+阅读 · 2023年6月15日

Active Representation Learning for General Task Space with Applications in Robotics

Arxiv

0+阅读 · 2023年6月15日

Offline Multi-Agent Reinforcement Learning with Coupled Value Factorization

Arxiv

0+阅读 · 2023年6月15日

Langevin Thompson Sampling with Logarithmic Communication: Bandits and Reinforcement Learning

Arxiv

0+阅读 · 2023年6月15日

Provably Efficient Offline Reinforcement Learning with Perturbed Data Sources

Arxiv

0+阅读 · 2023年6月14日

Dynamic Interval Restrictions on Action Spaces in Deep Reinforcement Learning for Obstacle Avoidance

Arxiv

0+阅读 · 2023年6月13日

VIP会员

文章信息

相关主题

最新内容

纵深侦察：大规模作战行动中远程侦察与监视之迫切需求

纵深侦察：大规模作战行动中远程侦察与监视之迫切需求

专知会员服务

0+阅读 · 26分钟前

共享认知，分布式研判：复杂行动中的美国空军指挥控制（万字长文）

共享认知，分布式研判：复杂行动中的美国空军指挥控制（万字长文）

专知会员服务

0+阅读 · 56分钟前

《无人机对海面作战影响评估》

《无人机对海面作战影响评估》

专知会员服务

11+阅读 · 7月21日

《可损耗无人系统规模化应用对美国军事转型的战略影响（2022-2030）》2026年270页

《可损耗无人系统规模化应用对美国军事转型的战略影响（2022-2030）》2026年270页

专知会员服务

10+阅读 · 7月21日

博士论文 | 后训练如何损害大模型生成多样性？SimpleStrat与Stylus

博士论文 | 后训练如何损害大模型生成多样性？SimpleStrat与Stylus

专知会员服务

4+阅读 · 7月21日

综述 | 面向5G/6G网络的LLM智能体AI：架构、协议与标准化

综述 | 面向5G/6G网络的LLM智能体AI：架构、协议与标准化

专知会员服务

6+阅读 · 7月21日

五角大楼新设无人机办公室（DRPM-UxS）将如何重塑美国无人系统格局（附美国防部设立备忘录）

五角大楼新设无人机办公室（DRPM-UxS）将如何重塑美国无人系统格局（附美国防部设立备忘录）

专知会员服务

8+阅读 · 7月21日

印度精确打击与指挥架构的断层

印度精确打击与指挥架构的断层

专知会员服务

6+阅读 · 7月20日

《NASA喷气推进实验室：高耐久轻质常驻空观测系统（HELIOS）》429页

《NASA喷气推进实验室：高耐久轻质常驻空观测系统（HELIOS）》429页

专知会员服务

8+阅读 · 7月20日

美空军AI完成F-16战斗机自主空战历史性试飞

美空军AI完成F-16战斗机自主空战历史性试飞

专知会员服务

6+阅读 · 7月20日

《美政府问责局——武器系统年度评估（2026年）：强制要求成熟技术或可推动转向快速交付》249页

《美政府问责局——武器系统年度评估（2026年）：强制要求成熟技术或可推动转向快速交付》249页

专知会员服务

9+阅读 · 7月20日

《美国陆军：通过弹性分布式模型库实现自适应AI优势》

《美国陆军：通过弹性分布式模型库实现自适应AI优势》

专知会员服务

8+阅读 · 7月20日

博士论文 | 理解与改进大语言模型推理：从反转诅咒到连续思维链

博士论文 | 理解与改进大语言模型推理：从反转诅咒到连续思维链

专知会员服务

10+阅读 · 7月20日

综述 | 终身视觉表征：持续自监督学习CSSL系统综述

综述 | 终身视觉表征：持续自监督学习CSSL系统综述

专知会员服务

10+阅读 · 7月20日

深入Project Maven：为何人工智能在战场上依然失灵

深入Project Maven：为何人工智能在战场上依然失灵

专知会员服务

15+阅读 · 7月19日

相关VIP内容

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

专知会员服务

247+阅读 · 2019年10月21日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

37+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

61+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

60+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

164+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

共享认知，分布式研判：复杂行动中的美国空军指挥控制（万字长文）

《可损耗无人系统规模化应用对美国军事转型的战略影响（2022-2030）》2026年270页

纵深侦察：大规模作战行动中远程侦察与监视之迫切需求

《无人机对海面作战影响评估》

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Hierarchical Imitation - Reinforcement Learning

Hierarchical Imitation - Reinforcement Learning

CreateAMind

19+阅读 · 2018年5月25日

Capsule Networks解析

Capsule Networks解析

机器学习研究会

11+阅读 · 2017年11月12日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Robotic Packaging Optimization with Reinforcement Learning

Arxiv

0+阅读 · 2023年6月16日

Semi-Offline Reinforcement Learning for Optimized Text Generation

Arxiv

0+阅读 · 2023年6月16日

Generalizable Resource Scaling of 5G Slices using Constrained Reinforcement Learning

Arxiv

0+阅读 · 2023年6月15日

Neuroevolution is a Competitive Alternative to Reinforcement Learning for Skill Discovery

Arxiv

0+阅读 · 2023年6月15日

RLTP: Reinforcement Learning to Pace for Delayed Impression Modeling in Preloaded Ads

Arxiv

0+阅读 · 2023年6月15日

Active Representation Learning for General Task Space with Applications in Robotics

Arxiv

0+阅读 · 2023年6月15日

Offline Multi-Agent Reinforcement Learning with Coupled Value Factorization

Arxiv

0+阅读 · 2023年6月15日

Langevin Thompson Sampling with Logarithmic Communication: Bandits and Reinforcement Learning

Arxiv

0+阅读 · 2023年6月15日

Provably Efficient Offline Reinforcement Learning with Perturbed Data Sources

Arxiv

0+阅读 · 2023年6月14日

Dynamic Interval Restrictions on Action Spaces in Deep Reinforcement Learning for Obstacle Avoidance

Arxiv

0+阅读 · 2023年6月13日

相关基金

疏肝法对情绪调节不良MCI患者工作记忆影响的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

离子取代与晶格调控对YAG:Ce3+荧光粉发光性质的影响

国家自然科学基金

0+阅读 · 2014年12月31日

平流层突然增温期间电离层赤道异常区周期性变化研究

国家自然科学基金

0+阅读 · 2014年12月31日

FRP加固钢筋混凝土柱受压性能的尺寸效应研究及工程应用

国家自然科学基金

0+阅读 · 2013年12月31日

慢性应激对肿瘤术后复发和转移的影响及机制探讨

国家自然科学基金

0+阅读 · 2013年12月31日

MDSCs在动脉粥样硬化中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

Pharicin B稳定维甲酸受体的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

随机变分不等式

国家自然科学基金

0+阅读 · 2011年12月31日

基于HM耦合效应的高渗透水压水工隧洞承载机理研究

国家自然科学基金

0+阅读 · 2011年12月31日

离子类溶质在土中迁移过程的耦合效应仿真分析

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员