Surrogate rewards for linear temporal logic (LTL) objectives are commonly utilized in planning problems for LTL objectives. In a widely-adopted surrogate reward approach, two discount factors are used to ensure that the expected return approximates the satisfaction probability of the LTL objective. The expected return then can be estimated by methods using the Bellman updates such as reinforcement learning. However, the uniqueness of the solution to the Bellman equation with two discount factors has not been explicitly discussed. We demonstrate with an example that when one of the discount factors is set to one, as allowed in many previous works, the Bellman equation may have multiple solutions, leading to inaccurate evaluation of the expected return. We then propose a condition for the Bellman equation to have the expected return as the unique solution, requiring the solutions for states inside a rejecting bottom strongly connected component (BSCC) to be 0. We prove this condition is sufficient by showing that the solutions for the states with discounting can be separated from those for the states without discounting under this condition
翻译:针对线性时序逻辑(LTL)目标的替代奖励函数,在LTL目标的规划问题中得到了广泛应用。一种常用的替代奖励方法通过引入两个折扣因子,确保期望回报能近似逼近LTL目标的满足概率。随后,该期望回报可通过基于Bellman更新的方法(如强化学习)进行估计。然而,含两个折扣因子的Bellman方程解的唯一性问题尚未得到明确讨论。我们通过一个示例表明:当一个折扣因子设为1时(如众多前期研究所允许的情况),Bellman方程可能存在多解,导致期望回报的评估不准确。为此,我们提出一个使期望回报成为Bellman方程唯一解的条件,要求拒绝型底强连通分量(BSCC)内状态对应的解必须为零。通过证明在该条件下,含折扣状态与无折扣状态的解可实现分离,我们验证了该条件的充分性。