We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance-difference and local-approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/ proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach.
翻译:我们研究了连续时间和空间环境下的强化学习(RL),针对无限时域、折扣目标函数且底层动态由随机微分方程驱动的情形。基于连续强化学习方法的最新进展,我们提出了占用时间(特别是针对折扣目标函数)的概念,并展示了如何有效利用它推导性能差异公式和局部近似公式。我们进一步扩展这些结果,展示了它们在策略梯度(PG)、信任区域策略优化/近端策略优化(TRPO/PPO)方法中的应用——这些方法在离散RL中已是成熟且强大的工具,但在连续RL中尚未充分发展。通过数值实验,我们验证了所提方法的有效性和优势。