Autonomous agents operating in real-world scenarios frequently encounter uncertainty and make decisions based on incomplete information. Planning under uncertainty can be mathematically formalized using partially observable Markov decision processes (POMDPs). However, finding an optimal plan for POMDPs can be computationally expensive and is feasible only for small tasks. In recent years, approximate algorithms, such as tree search and sample-based methodologies, have emerged as state-of-the-art POMDP solvers for larger problems. Despite their effectiveness, these algorithms offer only probabilistic and often asymptotic guarantees toward the optimal solution due to their dependence on sampling. To address these limitations, we derive a deterministic relationship between a simplified solution that is easier to obtain and the theoretically optimal one. First, we derive bounds for selecting a subset of the observations to branch from while computing a complete belief at each posterior node. Then, since a complete belief update may be computationally demanding, we extend the bounds to support reduction of both the state and the observation spaces. We demonstrate how our guarantees can be integrated with existing state-of-the-art solvers that sample a subset of states and observations. As a result, the returned solution holds deterministic bounds relative to the optimal policy. Lastly, we substantiate our findings with supporting experimental results.
翻译:在现实场景中运行的自主智能体经常面临不确定性,并基于不完整信息做出决策。不确定性下的规划可以通过部分可观测马尔可夫决策过程(POMDP)进行数学形式化。然而,为POMDP寻找最优规划的计算代价高昂,仅适用于小规模任务。近年来,树搜索和基于样本的方法等近似算法已成为处理更大规模问题的先进POMDP求解器。尽管这些算法效果显著,但由于依赖采样,它们仅能提供面向最优解的概率性、且往往是渐近性的保证。为解决这些局限,我们推导出易于获取的简化解与理论最优解之间的确定性关系。首先,我们推导了在计算每个后验节点的完整信念时,选择用于分支的观测子集的边界。接着,考虑到完整信念更新可能计算负担沉重,我们将边界扩展以支持状态空间和观测空间的缩减。我们展示了如何将保证与现有先进求解器(这些求解器采样状态和观测的子集)相结合。因此,返回的解相对于最优策略具有确定性边界。最后,我们通过支持性实验结果验证了我们的发现。