In robust Markov decision processes (MDPs), the uncertainty in the transition kernel is addressed by finding a policy that optimizes the worst-case performance over an uncertainty set of MDPs. While much of the literature has focused on discounted MDPs, robust average-reward MDPs remain largely unexplored. In this paper, we focus on robust average-reward MDPs, where the goal is to find a policy that optimizes the worst-case average reward over an uncertainty set. We first take an approach that approximates average-reward MDPs using discounted MDPs. We prove that the robust discounted value function converges to the robust average-reward as the discount factor $\gamma$ goes to $1$, and moreover, when $\gamma$ is large, any optimal policy of the robust discounted MDP is also an optimal policy of the robust average-reward. We further design a robust dynamic programming approach, and theoretically characterize its convergence to the optimum. Then, we investigate robust average-reward MDPs directly without using discounted MDPs as an intermediate step. We derive the robust Bellman equation for robust average-reward MDPs, prove that the optimal policy can be derived from its solution, and further design a robust relative value iteration algorithm that provably finds its solution, or equivalently, the optimal robust policy.
翻译:在鲁棒马尔可夫决策过程中,转移核的不确定性通过寻找一个在马尔可夫决策过程不确定性集上优化最坏情况性能的策略来处理。尽管大量文献关注于折扣马尔可夫决策过程,鲁棒平均奖励马尔可夫决策过程仍未得到充分探索。本文重点研究鲁棒平均奖励马尔可夫决策过程,其目标是在不确定性集上寻找一个能优化最坏情况平均奖励的策略。我们首先采用一种使用折扣马尔可夫决策过程近似平均奖励马尔可夫决策过程的方法。我们证明,当折扣因子$\gamma$趋近于1时,鲁棒折扣值函数收敛至鲁棒平均奖励;此外,当$\gamma$较大时,鲁棒折扣马尔可夫决策过程的任一最优策略也是鲁棒平均奖励马尔可夫决策过程的最优策略。我们进一步设计了一种鲁棒动态规划方法,并从理论上刻画了其收敛到最优解的特性。随后,我们直接研究鲁棒平均奖励马尔可夫决策过程,而不将折扣马尔可夫决策过程作为中间步骤。我们推导出鲁棒平均奖励马尔可夫决策过程的鲁棒贝尔曼方程,证明其解可导出最优策略,并进一步设计了一种鲁棒相对值迭代算法,该算法可证明地找到其解,即等价于找到最优鲁棒策略。