The problem of constrained reinforcement learning (CRL) holds significant importance as it provides a framework for addressing critical safety satisfaction concerns in the field of reinforcement learning (RL). However, with the introduction of constraint satisfaction, the current CRL methods necessitate the utilization of second-order optimization or primal-dual frameworks with additional Lagrangian multipliers, resulting in increased complexity and inefficiency during implementation. To address these issues, we propose a novel first-order feasible method named Constrained Proximal Policy Optimization (CPPO). By treating the CRL problem as a probabilistic inference problem, our approach integrates the Expectation-Maximization framework to solve it through two steps: 1) calculating the optimal policy distribution within the feasible region (E-step), and 2) conducting a first-order update to adjust the current policy towards the optimal policy obtained in the E-step (M-step). We establish the relationship between the probability ratios and KL divergence to convert the E-step into a convex optimization problem. Furthermore, we develop an iterative heuristic algorithm from a geometric perspective to solve this problem. Additionally, we introduce a conservative update mechanism to overcome the constraint violation issue that occurs in the existing feasible region method. Empirical evaluations conducted in complex and uncertain environments validate the effectiveness of our proposed method, as it performs at least as well as other baselines.
翻译:受限强化学习(CRL)问题具有重要意义,因为它为解决强化学习(RL)领域中的关键安全性满足问题提供了框架。然而,随着约束满足的引入,当前CRL方法需利用二阶优化或带有额外拉格朗日乘子的原始-对偶框架,导致实现复杂度增加和效率降低。为解决这些问题,我们提出了一种新颖的一阶可行方法,命名为受限近端策略优化(CPPO)。通过将CRL问题视为概率推断问题,我们的方法集成期望最大化框架,通过两步求解:1)在可行区域内计算最优策略分布(E步),2)执行一阶更新以将当前策略调整至E步得到的最优策略(M步)。我们建立了概率比率与KL散度之间的关系,将E步转化为凸优化问题。此外,我们从几何角度开发了一种迭代启发式算法来求解该问题。同时,引入保守更新机制以克服现有可行区域方法中出现的约束违反问题。在复杂和不确定环境中进行的实证评估验证了所提方法的有效性,其性能至少不逊于其他基线方法。