Offline goal-conditioned reinforcement learning (GCRL) learns goal-reaching behaviors from static datasets, but accurate value estimation remains challenging under limited state-action coverage. Existing physics-informed approaches address this by imposing pointwise distance-like geometric constraints derived from Hamilton--Jacobi--Bellman (HJB) optimality principles, often through first-order partial differential equations such as the Eikonal equation. However, enforcing local consistency through explicit differential structure can become unstable in complex, high-dimensional environments. Our key insight is to instead reinterpret distance-like constraints as an expectation over a local spatial measure. By aggregating constraints over this measure rather than evaluating them pointwise, the objective acts as a spatial mollifier, inducing distance-like value geometry without requiring expensive differential operators. We refer to this as Mollified Value Learning (MVL). Experiments across navigation and high-dimensional robotic manipulation tasks show that MVL learns structured, value representations, improving goal-reaching performance, when used with implicit value representation learning methods. Open-source codes are available at https://github.com/HrishikeshVish/MVL.
翻译:离线目标条件强化学习(GCRL)从静态数据集中学习目标导向行为,但在有限状态-动作覆盖下,准确的值估计仍具挑战性。现有物理信息方法通过施加源于汉密尔顿-雅可比-贝尔曼(HJB)最优性原理的点态距离类几何约束(通常借助一阶偏微分方程如程函方程)来解决这一问题。然而,在复杂高维环境中,通过显式微分结构施加局部一致性可能变得不稳定。我们的核心见解在于将距离类约束重新解释为局部空间测度上的期望。通过在该测度上聚合约束而非逐点评估,目标函数充当空间平滑算子,无需昂贵微分算子即可诱导距离类值几何结构。我们将其称为平滑值学习(MVL)。在导航和高维机器人操作任务上的实验表明,当与隐式值表征学习方法结合使用时,MVL可学习结构化的值表征,提升目标到达性能。开源代码见https://github.com/HrishikeshVish/MVL。