In the dynamic and uncertain environments where reinforcement learning (RL) operates, risk management becomes a crucial factor in ensuring reliable decision-making. Traditional RL approaches, while effective in reward optimization, often overlook the landscape of potential risks. In response, this paper pioneers the integration of Optimal Transport (OT) theory with RL to create a risk-aware framework. Our approach modifies the objective function, ensuring that the resulting policy not only maximizes expected rewards but also respects risk constraints dictated by OT distances between state visitation distributions and the desired risk profiles. By leveraging the mathematical precision of OT, we offer a formulation that elevates risk considerations alongside conventional RL objectives. Our contributions are substantiated with a series of theorems, mapping the relationships between risk distributions, optimal value functions, and policy behaviors. Through the lens of OT, this work illuminates a promising direction for RL, ensuring a balanced fusion of reward pursuit and risk awareness.
翻译:在强化学习运行的动态不确定环境中,风险管理成为确保可靠决策的关键因素。传统强化学习方法虽在奖励优化方面表现有效,却常忽视潜在风险的全局图景。为此,本文开创性地将最优传输理论与强化学习相结合,构建了风险感知框架。该方法通过修改目标函数,确保所得策略不仅能最大化期望奖励,还能满足由状态访问分布与期望风险分布之间的最优传输距离所界定的风险约束。借助最优传输理论的数学精确性,我们提出了一种将风险考量与传统强化学习目标并重的数学形式化表达。通过一系列定理的证明,本文阐明了风险分布、最优值函数与策略行为之间的映射关系。借助最优传输理论的视角,本研究为强化学习开辟了富有前景的方向,实现了奖励追求与风险意识的均衡融合。