This paper studies the sample-efficiency of learning in Partially Observable Markov Decision Processes (POMDPs), a challenging problem in reinforcement learning that is known to be exponentially hard in the worst-case. Motivated by real-world settings such as loading in game playing, we propose an enhanced feedback model called ``multiple observations in hindsight'', where after each episode of interaction with the POMDP, the learner may collect multiple additional observations emitted from the encountered latent states, but may not observe the latent states themselves. We show that sample-efficient learning under this feedback model is possible for two new subclasses of POMDPs: \emph{multi-observation revealing POMDPs} and \emph{distinguishable POMDPs}. Both subclasses generalize and substantially relax \emph{revealing POMDPs} -- a widely studied subclass for which sample-efficient learning is possible under standard trajectory feedback. Notably, distinguishable POMDPs only require the emission distributions from different latent states to be \emph{different} instead of \emph{linearly independent} as required in revealing POMDPs.
翻译:本文研究了部分可观测马尔可夫决策过程(POMDP)中学习的样本效率问题,这是强化学习中一个最坏情况下已知呈指数级难度的挑战性问题。受游戏加载等现实场景启发,我们提出一种名为“事后多重观测”的增强反馈模型:在与POMDP进行每轮交互后,学习者可以收集来自所遇隐状态的多个额外观测,但无法直接观测隐状态本身。我们证明,在该反馈模型下,对于两类新的POMDP子类——即多观测揭示型POMDP和可区分型POMDP——可以实现样本高效学习。这两个子类均推广并实质性地放宽了揭示型POMDP(一个在标准轨迹反馈下可实现样本高效学习的广泛研究子类)。值得注意的是,可区分型POMDP仅要求不同隐状态的发射分布互异,而非如揭示型POMDP要求的线性独立。