We consider the reinforcement learning problem for partially observed Markov decision processes (POMDPs) with large or even countably infinite state spaces, where the controller has access to only noisy observations of the underlying controlled Markov chain. We consider a natural actor-critic method that employs a finite internal memory for policy parameterization, and a multi-step temporal difference learning algorithm for policy evaluation. We establish, to the best of our knowledge, the first non-asymptotic global convergence of actor-critic methods for partially observed systems under function approximation. In particular, in addition to the function approximation and statistical errors that also arise in MDPs, we explicitly characterize the error due to the use of finite-state controllers. This additional error is stated in terms of the total variation distance between the traditional belief state in POMDPs and the posterior distribution of the hidden state when using a finite-state controller. Further, we show that this error can be made small in the case of sliding-block controllers by using larger block sizes.
翻译:我们考虑状态空间庞大甚至可数无限的部分可观测马尔可夫决策过程(POMDPs)中的强化学习问题,其中控制器仅能访问受控马尔可夫链的含噪观测值。我们采用一种自然演员-评论家方法,该方法使用有限内部记忆进行策略参数化,并采用多步时序差分学习算法进行策略评估。据我们所知,这是首次在函数逼近条件下,为部分可观测系统建立演员-评论家方法的非渐近全局收敛性。特别地,除了马尔可夫决策过程(MDPs)中出现的函数逼近误差和统计误差外,我们明确刻画了使用有限状态控制器所带来的额外误差。该额外误差通过POMDPs中传统信念状态与使用有限状态控制器时隐状态后验分布之间的全变差距离进行表述。进一步分析表明,对于滑动块控制器,通过增大块尺寸可有效降低该误差。