Reinforcement learning (RL) has become the de facto method for achieving locomotion on humanoid robots in practice, yet stability analysis of the corresponding control policies is lacking. Recent work has attempted to merge control theoretic ideas with reinforcement learning through control guided learning. A notable example of this is the use of a control Lyapunov function (CLF) to synthesize the reinforcement learning rewards, a technique known as CLF-RL, which has shown practical success. This paper investigates the stability properties of optimal controllers using CLF-RL with the goal of bridging experimentally observed stability with theoretical guarantees. The RL problem is viewed as an optimal control problem and exponential stability is proven in both continuous and discrete time using both core CLF reward terms and the additional terms used in practice. The theoretical bounds are numerically verified on systems such as the double integrator and cart-pole. Finally, the CLF guided rewards are implemented for a walking humanoid robot to generate stable periodic orbits.


翻译:强化学习(RL)已成为实现人形机器人实际运动的事实标准方法,但相应控制策略的稳定性分析仍然缺乏。近期研究尝试通过控制引导学习将控制理论思想与强化学习相结合,其中利用控制李雅普诺夫函数(CLF)综合强化学习奖励的技术(即CLF-RL)是典型成功案例。本文旨在研究采用CLF-RL的最优控制器的稳定性特性,以期在实验观测稳定性与理论保证之间建立联系。本文将强化学习问题视为最优控制问题,并证明在连续时间和离散时间下,使用核心CLF奖励项及实际应用中的附加项均可实现指数稳定性。通过双积分器和车杆系统等案例对理论边界进行了数值验证。最后,将CLF引导奖励应用于行走人形机器人以生成稳定周期轨道。

0
下载
关闭预览

相关内容

《机器人强化学习技术进展》34页
专知会员服务
40+阅读 · 2025年7月16日
《用于水下目标定位的平台便携式强化学习方法》
专知会员服务
28+阅读 · 2024年1月2日
基于人工反馈的强化学习综述
专知会员服务
66+阅读 · 2023年12月25日
基于模型的强化学习综述
专知会员服务
150+阅读 · 2022年7月13日
专知会员服务
101+阅读 · 2021年7月11日
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
强化学习精品书籍
平均机器
26+阅读 · 2019年1月2日
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
5+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
9+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
11+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员