稳定Q梯度场以实现Actor-Critic中的策略平滑性 (Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic) - 专知论文

会员服务 ·

0

平滑 · 平滑性 · 梯度 · 评论员 · 梯度场 ·

Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic

翻译：稳定Q梯度场以实现Actor-Critic中的策略平滑性

Jeong Woon Lee,Kyoleen Kwak,Daeho Kim,Hyoseok Hwang

Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physical deployment. Current approaches attempt to enforce smoothness by directly regularizing the policy's output. We argue that this approach treats the symptom rather than the cause. In this work, we theoretically establish that policy non-smoothness is fundamentally governed by the differential geometry of the critic. By applying implicit differentiation to the actor-critic objective, we prove that the sensitivity of the optimal policy is bounded by the ratio of the Q-function's mixed-partial derivative (noise sensitivity) to its action-space curvature (signal distinctness). To empirically validate this theoretical insight, we introduce PAVE (Policy-Aware Value-field Equalization), a critic-centric regularization framework that treats the critic as a scalar field and stabilizes its induced action-gradient field. PAVE rectifies the learning signal by minimizing the Q-gradient volatility while preserving local curvature. Experimental results demonstrate that PAVE achieves smoothness and robustness comparable to policy-side smoothness regularization methods, while maintaining competitive task performance, without modifying the actor.

翻译：通过连续型Actor-Critic方法学习得到的策略常表现出不稳定、高频的振荡特性，使其难以部署于物理系统。现有方法试图通过直接正则化策略输出来强制实现平滑性。我们认为这种方法仅处理了表象而非根本原因。本文从理论上证明，策略的非平滑性本质上由评论家函数的微分几何特性所决定。通过对Actor-Critic目标函数应用隐函数微分法，我们证明了最优策略的敏感度受限于Q函数混合偏导数（噪声敏感度）与其动作空间曲率（信号区分度）的比值。为实证验证这一理论见解，我们提出了PAVE（策略感知价值场均衡化）——一种以评论家为核心的正则化框架，该方法将评论家视为标量场并稳定其诱导的动作梯度场。PAVE通过最小化Q梯度波动性同时保持局部曲率来修正学习信号。实验结果表明，在不修改行动者网络的前提下，PAVE实现的平滑性与鲁棒性可与策略侧平滑正则化方法相媲美，同时保持具有竞争力的任务性能。

0

相关内容

【斯坦福博士论文】高精度操控的策略学习前沿研究

【斯坦福博士论文】高精度操控的策略学习前沿研究

专知会员服务

22+阅读 · 2025年3月30日

【牛津大学博士论文】不确定性量化与因果考量在非策略决策制定中的应用

【牛津大学博士论文】不确定性量化与因果考量在非策略决策制定中的应用

专知会员服务

20+阅读 · 2025年2月24日

【MIT博士论文】非线性优化在机器学习应用中的平滑性与自适应性

【MIT博士论文】非线性优化在机器学习应用中的平滑性与自适应性

专知会员服务

27+阅读 · 2024年8月27日

非平稳过程异常监测方法：综述与展望

非平稳过程异常监测方法：综述与展望

专知会员服务

23+阅读 · 2024年7月16日

【ICML2022】鲁棒强化学习的策略梯度法

【ICML2022】鲁棒强化学习的策略梯度法

专知会员服务

38+阅读 · 2022年5月21日

【NeurIPS2021】Spatial Ensemble：一种新颖的用于学生-老师框架的模型平滑机制

【NeurIPS2021】Spatial Ensemble：一种新颖的用于学生-老师框架的模型平滑机制

专知会员服务

18+阅读 · 2021年11月8日

【NeurIPS 2021】设置多智能体策略梯度的方差

【NeurIPS 2021】设置多智能体策略梯度的方差

专知会员服务

21+阅读 · 2021年10月24日

通过条件梯度进行结构化机器学习训练，50页ppt与视频

通过条件梯度进行结构化机器学习训练，50页ppt与视频

专知会员服务

13+阅读 · 2021年2月25日

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

专知会员服务

17+阅读 · 2020年7月14日

策略梯度方法的算子视图，An operator view of policy gradient methods

策略梯度方法的算子视图，An operator view of policy gradient methods

专知会员服务

11+阅读 · 2020年6月23日

【佐治亚理工博士论文】基于策略智能体和有限反馈的序列决策，211页pdf

【佐治亚理工博士论文】基于策略智能体和有限反馈的序列决策，211页pdf

专知

38+阅读 · 2023年4月13日

《基于近端策略优化(PPO)算法的制导弹体控制行为学习》美国陆军2022最新27页技术报告

《基于近端策略优化(PPO)算法的制导弹体控制行为学习》美国陆军2022最新27页技术报告

专知

13+阅读 · 2022年11月25日

推荐！《不确定性下的作战决策：推理、序贯和对抗性方法》美国空军293页博士论文，含代码

推荐！《不确定性下的作战决策：推理、序贯和对抗性方法》美国空军293页博士论文，含代码

专知

47+阅读 · 2022年11月16日

【254页博士论文】《动态多目标环境中基于深度强化学习的智能决策方案》

【254页博士论文】《动态多目标环境中基于深度强化学习的智能决策方案》

专知

32+阅读 · 2022年10月17日

【干货书】深度不确定性条件下的决策:理论到实践，408页pdf

【干货书】深度不确定性条件下的决策:理论到实践，408页pdf

专知

17+阅读 · 2021年1月18日

Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证

Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证

深度强化学习实验室

19+阅读 · 2020年8月11日

强化学习开篇：Q-Learning原理详解

强化学习开篇：Q-Learning原理详解

AINLP

37+阅读 · 2020年7月28日

「PPT」深度学习中的不确定性估计

「PPT」深度学习中的不确定性估计

专知

27+阅读 · 2019年7月20日

深度强化学习首次在无监督视频摘要生成问题中的应用：实现state-of-the-art效果

深度强化学习首次在无监督视频摘要生成问题中的应用：实现state-of-the-art效果

专知

26+阅读 · 2018年1月21日

【AI唠科】Focal Loss：助大神何凯明获得ICCV最佳学生论文，究竟有什么功？|兼谈目标检测发展历程

【AI唠科】Focal Loss：助大神何凯明获得ICCV最佳学生论文，究竟有什么功？|兼谈目标检测发展历程

中国科学院自动化研究所

10+阅读 · 2017年11月16日

针对大规模环境下复杂任务的策略搜索强化学习方法研究

国家自然科学基金

42+阅读 · 2015年12月31日

基于滑模技术的线性参数变化系统的容错控制方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

高维回归模型的预测稳定性研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于犹豫模糊语言信息的定性决策理论与方法

国家自然科学基金

2+阅读 · 2015年12月31日

一类不确定非线性大系统的非光滑分散控制研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于扩展状态观测器的不确定分数阶系统镇定设计

国家自然科学基金

0+阅读 · 2015年12月31日

非线性不确定系统的齐次控制理论及应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

结构振动的非光滑控制方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

多域网络安全的异构策略语义形态与验证机制

国家自然科学基金

0+阅读 · 2014年12月31日

不确定非凸规划的稳健全局优化方法的研究

国家自然科学基金

1+阅读 · 2014年12月31日

Certified Gradient-Based Contact-Rich Manipulation via Smoothing-Error Reachable Tubes

Arxiv

0+阅读 · 2月10日

Direct Soft-Policy Sampling via Langevin Dynamics

Arxiv

0+阅读 · 2月8日

Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration

Arxiv

0+阅读 · 2月8日

From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

Arxiv

0+阅读 · 2月4日

Performative Policy Gradient: Optimality in Performative Reinforcement Learning

Arxiv

0+阅读 · 2月2日

Reusing Trajectories in Policy Gradients Enables Fast Convergence

Arxiv

0+阅读 · 2月2日

Model-free policy gradient for discrete-time mean-field control

Arxiv

0+阅读 · 1月27日

Scaling Effects and Uncertainty Quantification in Neural Actor Critic Algorithms

Arxiv

0+阅读 · 1月25日

Stabilizing Policy Gradient Methods via Reward Profiling

Arxiv

0+阅读 · 1月24日

Off Policy Lyapunov Stability in Reinforcement Learning

Arxiv

0+阅读 · 1月16日

VIP会员

文章信息

相关主题

相关VIP内容

【斯坦福博士论文】高精度操控的策略学习前沿研究

【斯坦福博士论文】高精度操控的策略学习前沿研究

专知会员服务

22+阅读 · 2025年3月30日

【牛津大学博士论文】不确定性量化与因果考量在非策略决策制定中的应用

【牛津大学博士论文】不确定性量化与因果考量在非策略决策制定中的应用

专知会员服务

20+阅读 · 2025年2月24日

【MIT博士论文】非线性优化在机器学习应用中的平滑性与自适应性

【MIT博士论文】非线性优化在机器学习应用中的平滑性与自适应性

专知会员服务

27+阅读 · 2024年8月27日

非平稳过程异常监测方法：综述与展望

非平稳过程异常监测方法：综述与展望

专知会员服务

23+阅读 · 2024年7月16日

【ICML2022】鲁棒强化学习的策略梯度法

【ICML2022】鲁棒强化学习的策略梯度法

专知会员服务

38+阅读 · 2022年5月21日

【NeurIPS2021】Spatial Ensemble：一种新颖的用于学生-老师框架的模型平滑机制

【NeurIPS2021】Spatial Ensemble：一种新颖的用于学生-老师框架的模型平滑机制

专知会员服务

18+阅读 · 2021年11月8日

【NeurIPS 2021】设置多智能体策略梯度的方差

【NeurIPS 2021】设置多智能体策略梯度的方差

专知会员服务

21+阅读 · 2021年10月24日

通过条件梯度进行结构化机器学习训练，50页ppt与视频

通过条件梯度进行结构化机器学习训练，50页ppt与视频

专知会员服务

13+阅读 · 2021年2月25日

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

【ICML2020-伯克利】稳定非策略强化学习的表示，Representations for Stable Off-Policy Reinforcement Learning

专知会员服务

17+阅读 · 2020年7月14日

策略梯度方法的算子视图，An operator view of policy gradient methods

策略梯度方法的算子视图，An operator view of policy gradient methods

专知会员服务

11+阅读 · 2020年6月23日

热门VIP内容

开通专知VIP会员享更多权益服务

智能体记忆深度剖析：评价指标与系统局限性的分类体系及实证分析

《可信人工智能赋能系统的支柱》

【CMU博士论文】可靠轨迹预测的分层基石：数据、评估与方法

人工智能赋能边缘与自主系统：美陆军现代化进程聚焦威胁探测与战术边缘情报

相关资讯

【佐治亚理工博士论文】基于策略智能体和有限反馈的序列决策，211页pdf

【佐治亚理工博士论文】基于策略智能体和有限反馈的序列决策，211页pdf

专知

38+阅读 · 2023年4月13日

《基于近端策略优化(PPO)算法的制导弹体控制行为学习》美国陆军2022最新27页技术报告

《基于近端策略优化(PPO)算法的制导弹体控制行为学习》美国陆军2022最新27页技术报告

专知

13+阅读 · 2022年11月25日

推荐！《不确定性下的作战决策：推理、序贯和对抗性方法》美国空军293页博士论文，含代码

推荐！《不确定性下的作战决策：推理、序贯和对抗性方法》美国空军293页博士论文，含代码

专知

47+阅读 · 2022年11月16日

【254页博士论文】《动态多目标环境中基于深度强化学习的智能决策方案》

【254页博士论文】《动态多目标环境中基于深度强化学习的智能决策方案》

专知

32+阅读 · 2022年10月17日

【干货书】深度不确定性条件下的决策:理论到实践，408页pdf

【干货书】深度不确定性条件下的决策:理论到实践，408页pdf

专知

17+阅读 · 2021年1月18日

Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证

Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证

深度强化学习实验室

19+阅读 · 2020年8月11日

强化学习开篇：Q-Learning原理详解

强化学习开篇：Q-Learning原理详解

AINLP

37+阅读 · 2020年7月28日

「PPT」深度学习中的不确定性估计

「PPT」深度学习中的不确定性估计

专知

27+阅读 · 2019年7月20日

深度强化学习首次在无监督视频摘要生成问题中的应用：实现state-of-the-art效果

深度强化学习首次在无监督视频摘要生成问题中的应用：实现state-of-the-art效果

专知

26+阅读 · 2018年1月21日

【AI唠科】Focal Loss：助大神何凯明获得ICCV最佳学生论文，究竟有什么功？|兼谈目标检测发展历程

【AI唠科】Focal Loss：助大神何凯明获得ICCV最佳学生论文，究竟有什么功？|兼谈目标检测发展历程

中国科学院自动化研究所

10+阅读 · 2017年11月16日

相关论文

Certified Gradient-Based Contact-Rich Manipulation via Smoothing-Error Reachable Tubes

Arxiv

0+阅读 · 2月10日

Direct Soft-Policy Sampling via Langevin Dynamics

Arxiv

0+阅读 · 2月8日

Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration

Arxiv

0+阅读 · 2月8日

From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

Arxiv

0+阅读 · 2月4日

Performative Policy Gradient: Optimality in Performative Reinforcement Learning

Arxiv

0+阅读 · 2月2日

Reusing Trajectories in Policy Gradients Enables Fast Convergence

Arxiv

0+阅读 · 2月2日

Model-free policy gradient for discrete-time mean-field control

Arxiv

0+阅读 · 1月27日

Scaling Effects and Uncertainty Quantification in Neural Actor Critic Algorithms

Arxiv

0+阅读 · 1月25日

Stabilizing Policy Gradient Methods via Reward Profiling

Arxiv

0+阅读 · 1月24日

Off Policy Lyapunov Stability in Reinforcement Learning

Arxiv

0+阅读 · 1月16日

相关基金

针对大规模环境下复杂任务的策略搜索强化学习方法研究

国家自然科学基金

42+阅读 · 2015年12月31日

基于滑模技术的线性参数变化系统的容错控制方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

高维回归模型的预测稳定性研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于犹豫模糊语言信息的定性决策理论与方法

国家自然科学基金

2+阅读 · 2015年12月31日

一类不确定非线性大系统的非光滑分散控制研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于扩展状态观测器的不确定分数阶系统镇定设计

国家自然科学基金

0+阅读 · 2015年12月31日

非线性不确定系统的齐次控制理论及应用研究

国家自然科学基金

0+阅读 · 2015年12月31日

结构振动的非光滑控制方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

多域网络安全的异构策略语义形态与验证机制

国家自然科学基金

0+阅读 · 2014年12月31日

不确定非凸规划的稳健全局优化方法的研究

国家自然科学基金

1+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员