On the Global Convergence of Risk-Averse Policy Gradient Methods with Expected Conditional Risk Measures - 专知论文

会员服务 ·

0

条件风险 · 控制器 · Performer · ENJOY · 有向 ·

2023 年 5 月 30 日

On the Global Convergence of Risk-Averse Policy Gradient Methods with Expected Conditional Risk Measures

翻译：关于风险厌恶策略梯度方法在期望条件风险测度下的全局收敛性研究

Xian Yu,Lei Ying

Risk-sensitive reinforcement learning (RL) has become a popular tool to control the risk of uncertain outcomes and ensure reliable performance in various sequential decision-making problems. While policy gradient methods have been developed for risk-sensitive RL, it remains unclear if these methods enjoy the same global convergence guarantees as in the risk-neutral case. In this paper, we consider a class of dynamic time-consistent risk measures, called Expected Conditional Risk Measures (ECRMs), and derive policy gradient updates for ECRM-based objective functions. Under both constrained direct parameterization and unconstrained softmax parameterization, we provide global convergence and iteration complexities of the corresponding risk-averse policy gradient algorithms. We further test risk-averse variants of REINFORCE and actor-critic algorithms to demonstrate the efficacy of our method and the importance of risk control.

翻译：风险敏感强化学习已成为控制不确定结果风险、确保各类序贯决策问题中可靠性能的流行工具。尽管已有针对风险敏感强化学习的策略梯度方法，但这些方法是否具有与风险中性情形相同的全局收敛保证仍不明确。本文考虑一类称为期望条件风险测度的动态时间一致风险测度，并推导了基于ECRM目标函数的策略梯度更新规则。在有约束直接参数化与无约束柔性最大参数化两种情形下，我们给出了相应风险厌恶策略梯度算法的全局收敛性及迭代复杂度。我们进一步测试了REINFORCE和演员-评论家算法的风险厌恶变体，以证明我们方法的有效性及风险控制的重要性。

0

相关内容

条件风险

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

南大《优化方法（Optimization Methods》课程，推荐！

南大《优化方法（Optimization Methods》课程，推荐！

专知会员服务

80+阅读 · 2022年4月3日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

52+阅读 · 2020年12月14日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

针刺、语言任务干预卒中后运动性失语的fMRI/ERP双模态脑网络效应机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

脂联素线粒体稳态调节在糖尿病缺血心肌保护中的作用及关键分子机制

国家自然科学基金

0+阅读 · 2014年12月31日

基于Metasurface的THz慢波器件研究

国家自然科学基金

0+阅读 · 2013年12月31日

老年人视觉方位、方向辨别能力衰退的神经机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

超短超强激光驱动的高亮度Betatron辐射光源

国家自然科学基金

1+阅读 · 2013年12月31日

网络环境下非线性时变随机系统的最优递推滤波研究

国家自然科学基金

0+阅读 · 2013年12月31日

中缝背核中Ca2+和L-、T-及N-型钙通道对睡眠-觉醒的调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

从头设计蛋白质DS119折叠机制的分子模拟研究

国家自然科学基金

0+阅读 · 2012年12月31日

具选择功能的分布式合作控制系统

国家自然科学基金

0+阅读 · 2011年12月31日

GATA-4、MEF2A表达调控与失重致心肌细胞凋亡的相关性研究

国家自然科学基金

0+阅读 · 2009年12月31日

Convergence Guarantees for Stochastic Subgradient Methods in Nonsmooth Nonconvex Optimization

Arxiv

0+阅读 · 2023年7月19日

Hierarchically Composing Level Generators for the Creation of Complex Structures

Arxiv

0+阅读 · 2023年7月19日

Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective

Arxiv

0+阅读 · 2023年7月18日

Stability and Generalization of Stochastic Optimization with Nonconvex and Nonsmooth Problems

Arxiv

0+阅读 · 2023年7月18日

An Alternative to Variance: Gini Deviation for Risk-averse Policy Gradient

Arxiv

0+阅读 · 2023年7月17日

Understanding Best Subset Selection: A Tale of Two C(omplex)ities

Arxiv

0+阅读 · 2023年7月17日

Robust empirical risk minimization via Newton's method

Arxiv

0+阅读 · 2023年7月17日

A subgradient method with constant step-size for $\ell_1$-composite optimization

Arxiv

0+阅读 · 2023年7月17日

Efficient numerical method for multi-term time-fractional diffusion equations with Caputo-Fabrizio derivatives

Arxiv

0+阅读 · 2023年7月16日

Identifiability Guarantees for Causal Disentanglement from Soft Interventions

Arxiv

0+阅读 · 2023年7月13日

VIP会员

文章信息

相关主题

最新内容

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

专知会员服务

2+阅读 · 6月16日

多模态代码智能综述：从视觉输入到可执行代码系统

多模态代码智能综述：从视觉输入到可执行代码系统

专知会员服务

1+阅读 · 6月16日

美国马六甲“三重网”概念：安全网、威慑网与杀伤网

美国马六甲“三重网”概念：安全网、威慑网与杀伤网

专知会员服务

4+阅读 · 6月16日

《面向导弹有效发射时机的监督机器学习方法：基于超视距空战仿真》

《面向导弹有效发射时机的监督机器学习方法：基于超视距空战仿真》

专知会员服务

3+阅读 · 6月16日

《通用大语言模型：无人机指挥与控制接口》最新40页

《通用大语言模型：无人机指挥与控制接口》最新40页

专知会员服务

13+阅读 · 6月16日

《通过小型无人机系统将情报能力“作战化”》

《通过小型无人机系统将情报能力“作战化”》

专知会员服务

4+阅读 · 6月16日

《神经安全型有人–无人协同：面向认知自适应作战能力的参考架构》

《神经安全型有人–无人协同：面向认知自适应作战能力的参考架构》

专知会员服务

8+阅读 · 6月16日

《在指挥链中通过多准则决策分析传达指挥官意图：空战实验》

《在指挥链中通过多准则决策分析传达指挥官意图：空战实验》

专知会员服务

20+阅读 · 6月15日

消耗优势：美军的“精确规模化”概念

消耗优势：美军的“精确规模化”概念

专知会员服务

8+阅读 · 6月15日

五角大楼的AI优先战略及其对现代战争的启示：来自与伊朗冲突的经验教训

五角大楼的AI优先战略及其对现代战争的启示：来自与伊朗冲突的经验教训

专知会员服务

9+阅读 · 6月15日

《网络空间兵棋推演：挑战、局限性与混合路径》报告

《网络空间兵棋推演：挑战、局限性与混合路径》报告

专知会员服务

9+阅读 · 6月15日

《离线语言支持系统：面向空战战术决策》

《离线语言支持系统：面向空战战术决策》

专知会员服务

9+阅读 · 6月15日

《以通信为中心的6G–LLM架构：面向可扩展的战术自主防御车辆网络》

《以通信为中心的6G–LLM架构：面向可扩展的战术自主防御车辆网络》

专知会员服务

8+阅读 · 6月15日

ICML 2026｜ECA：面向开放式图文生成的高效持续对齐

ICML 2026｜ECA：面向开放式图文生成的高效持续对齐

专知会员服务

6+阅读 · 6月14日

可信智能体AI综述：安全、鲁棒性、隐私与系统安全

可信智能体AI综述：安全、鲁棒性、隐私与系统安全

专知会员服务

6+阅读 · 6月14日

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

南大《优化方法（Optimization Methods》课程，推荐！

南大《优化方法（Optimization Methods》课程，推荐！

专知会员服务

80+阅读 · 2022年4月3日

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

INRIA最新「机器学习理论」新书，229页pdf原理性阐述机器学习

专知会员服务

69+阅读 · 2021年3月27日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

52+阅读 · 2020年12月14日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

多模态代码智能综述：从视觉输入到可执行代码系统

《面向导弹有效发射时机的监督机器学习方法：基于超视距空战仿真》

ICML 2026 | VOTP：用视频基础模型与最优传输，让离线偏好强化学习只需少量反馈

美国马六甲“三重网”概念：安全网、威慑网与杀伤网

相关资讯

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Convergence Guarantees for Stochastic Subgradient Methods in Nonsmooth Nonconvex Optimization

Arxiv

0+阅读 · 2023年7月19日

Hierarchically Composing Level Generators for the Creation of Complex Structures

Arxiv

0+阅读 · 2023年7月19日

Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective

Arxiv

0+阅读 · 2023年7月18日

Stability and Generalization of Stochastic Optimization with Nonconvex and Nonsmooth Problems

Arxiv

0+阅读 · 2023年7月18日

An Alternative to Variance: Gini Deviation for Risk-averse Policy Gradient

Arxiv

0+阅读 · 2023年7月17日

Understanding Best Subset Selection: A Tale of Two C(omplex)ities

Arxiv

0+阅读 · 2023年7月17日

Robust empirical risk minimization via Newton's method

Arxiv

0+阅读 · 2023年7月17日

A subgradient method with constant step-size for $\ell_1$-composite optimization

Arxiv

0+阅读 · 2023年7月17日

Efficient numerical method for multi-term time-fractional diffusion equations with Caputo-Fabrizio derivatives

Arxiv

0+阅读 · 2023年7月16日

Identifiability Guarantees for Causal Disentanglement from Soft Interventions

Arxiv

0+阅读 · 2023年7月13日

相关基金

针刺、语言任务干预卒中后运动性失语的fMRI/ERP双模态脑网络效应机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

脂联素线粒体稳态调节在糖尿病缺血心肌保护中的作用及关键分子机制

国家自然科学基金

0+阅读 · 2014年12月31日

基于Metasurface的THz慢波器件研究

国家自然科学基金

0+阅读 · 2013年12月31日

老年人视觉方位、方向辨别能力衰退的神经机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

超短超强激光驱动的高亮度Betatron辐射光源

国家自然科学基金

1+阅读 · 2013年12月31日

网络环境下非线性时变随机系统的最优递推滤波研究

国家自然科学基金

0+阅读 · 2013年12月31日

中缝背核中Ca2+和L-、T-及N-型钙通道对睡眠-觉醒的调控机制

国家自然科学基金

0+阅读 · 2012年12月31日

从头设计蛋白质DS119折叠机制的分子模拟研究

国家自然科学基金

0+阅读 · 2012年12月31日

具选择功能的分布式合作控制系统

国家自然科学基金

0+阅读 · 2011年12月31日

GATA-4、MEF2A表达调控与失重致心肌细胞凋亡的相关性研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员