We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance-difference and local-approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/ proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach.
翻译:我们研究连续时间和空间环境下的强化学习,考虑折扣目标下的无限时域问题,其中底层动力学由随机微分方程驱动。基于连续强化学习方法的最新进展,我们提出了一种占用时间(针对折扣目标)的概念,并展示了如何有效利用该概念推导性能差异公式和局部近似公式。进一步地,我们将这些结果扩展至策略梯度(PG)、信任域策略优化(TRPO)和近端策略优化(PPO)方法中,阐明其应用。这些方法在离散强化学习中已成为强大且熟悉的工具,但在连续强化学习中尚未充分发展。通过数值实验,我们证明了所提方法的有效性与优势。