Partially observable Markov decision processes (POMDPs) with continuous state and observation spaces have powerful flexibility for representing real-world decision and control problems but are notoriously difficult to solve. Recent online sampling-based algorithms that use observation likelihood weighting have shown unprecedented effectiveness in domains with continuous observation spaces. However there has been no formal theoretical justification for this technique. This work offers such a justification, proving that a simplified algorithm, partially observable weighted sparse sampling (POWSS), will estimate Q-values accurately with high probability and can be made to perform arbitrarily near the optimal solution by increasing computational power.
翻译:连续状态和观测空间的部分可观测马尔可夫决策过程(POMDP)在表示现实世界的决策与控制问题方面具有强大灵活性,但以其求解困难而著称。近年来,采用观测似然加权的在线采样算法在连续观测空间领域展现出前所未有的有效性。然而,该技术此前缺乏严格的理论证明。本研究为这一方法提供了理论依据,证明了一种简化算法——部分可观测加权稀疏采样(POWSS)——能够以高概率准确估计Q值,并且通过增加计算能力可使算法性能任意接近最优解。