Markov decision processes (MDPs) provide a standard framework for sequential decision making under uncertainty. However, transition probabilities in MDPs are often estimated from data and MDPs do not take data uncertainty into account. Robust Markov decision processes (RMDPs) address this shortcoming of MDPs by assigning to each transition an uncertainty set rather than a single probability value. The goal of solving RMDPs is then to find a policy which maximizes the worst-case performance over the uncertainty sets. In this work, we consider polytopic RMDPs in which all uncertainty sets are polytopes and study the problem of solving long-run average reward polytopic RMDPs. Our focus is on computational complexity aspects and efficient algorithms. We present a novel perspective on this problem and show that it can be reduced to solving long-run average reward turn-based stochastic games with finite state and action spaces. This reduction allows us to derive several important consequences that were hitherto not known to hold for polytopic RMDPs. First, we derive new computational complexity bounds for solving long-run average reward polytopic RMDPs, showing for the first time that the threshold decision problem for them is in NP coNP and that they admit a randomized algorithm with sub-exponential expected runtime. Second, we present Robust Polytopic Policy Iteration (RPPI), a novel policy iteration algorithm for solving long-run average reward polytopic RMDPs. Our experimental evaluation shows that RPPI is much more efficient in solving long-run average reward polytopic RMDPs compared to state-of-the-art methods based on value iteration.
翻译:马尔可夫决策过程(MDP)为不确定性环境下的序贯决策提供了标准框架。然而,MDP中的转移概率通常从数据中估计得出,且MDP并未考虑数据不确定性。鲁棒马尔可夫决策过程(RMDP)通过为每个转移分配一个不确定集而非单一概率值,弥补了MDP的这一缺陷。求解RMDP的目标是寻找一个策略,使其在不确定集上的最差性能最大化。本文考虑所有不确定集均为多面体的多面体RMDP,并研究求解长期平均奖励多面体RMDP的问题。我们重点关注计算复杂性方面及高效算法。我们对该问题提出了全新视角,证明其可归约为求解具有有限状态和动作空间的长期平均奖励回合制随机博弈。这一归约使我们能够推导出此前多面体RMDP尚未明确的若干重要结论。首先,我们为求解长期平均奖励多面体RMDP推导了新的计算复杂度上界,首次表明其阈值决策问题属于NP∩coNP,并且存在一个期望运行时间为次指数的随机化算法。其次,我们提出了鲁棒多面体策略迭代(RPPI)——一种用于求解长期平均奖励多面体RMDP的新型策略迭代算法。实验评估表明,与基于值迭代的现有最优方法相比,RPPI在求解长期平均奖励多面体RMDP时效率显著更高。