In real-world reinforcement learning problems, the state information is often only partially observable, which breaks the basic assumption in Markov decision processes, and thus, leads to inferior performances. Partially Observable Markov Decision Processes have been introduced to explicitly take the issue into account for learning, exploration, and planning, but presenting significant computational and statistical challenges. To address these difficulties, we exploit the representation view, which leads to a coherent design framework for a practically tractable reinforcement learning algorithm upon partial observations. We provide a theoretical analysis for justifying the statistical efficiency of the proposed algorithm. We also empirically demonstrate the proposed algorithm can surpass state-of-the-art performance with partial observations across various benchmarks, therefore, pushing reliable reinforcement learning towards more practical applications.
翻译:在现实世界的强化学习问题中,状态信息通常仅部分可观测,这破坏了马尔可夫决策过程的基本假设,从而导致性能不佳。部分可观测马尔可夫决策过程被引入以明确考虑学习、探索和规划中的这一问题,但带来了显著的计算和统计挑战。为克服这些困难,我们利用表示视角,为基于部分观测的实际可操作强化学习算法建立了一个连贯的设计框架。我们提供了理论分析以证明所提算法的统计效率。我们还通过实验表明,该算法在多种基准测试中能超越部分观测下的最先进性能,从而将可靠的强化学习推向更实际的应用场景。