Reinforcement learning (RL) has shown empirical success in various real world settings with complex models and large state-action spaces. The existing analytical results, however, typically focus on settings with a small number of state-actions or simple models such as linearly modeled state-action value functions. To derive RL policies that efficiently handle large state-action spaces with more general value functions, some recent works have considered nonlinear function approximation using kernel ridge regression. We propose $\pi$-KRVI, an optimistic modification of least-squares value iteration, when the state-action value function is represented by a reproducing kernel Hilbert space (RKHS). We prove the first order-optimal regret guarantees under a general setting. Our results show a significant polynomial in the number of episodes improvement over the state of the art. In particular, with highly non-smooth kernels (such as Neural Tangent kernel or some Mat\'ern kernels) the existing results lead to trivial (superlinear in the number of episodes) regret bounds. We show a sublinear regret bound that is order optimal in the case of Mat\'ern kernels where a lower bound on regret is known.
翻译:强化学习已在具有复杂模型和大状态-动作空间的各种实际场景中展现出经验成功。然而,现有理论分析通常聚焦于小规模状态-动作空间或简单模型(如线性建模的状态-动作值函数)的设置。为推导能高效处理大状态-动作空间及更一般值函数的强化学习策略,近期部分研究采用核岭回归进行非线性函数逼近。本文提出$\pi$-KRVI——一种基于最小二乘值迭代的乐观改进算法,其状态-动作值函数通过再生核希尔伯特空间表示。我们首次在一般设置下证明了最优阶遗憾界保证。研究结果表明,相较于现有最优方法,本文方法在回合数上实现了显著的多项式级提升。特别地,对于高度非光滑核(如神经正切核或某些马特恩核),现有方法仅能获得平凡(超线性于回合数)的遗憾界。我们则证明了次线性遗憾界,且在已知遗憾下界的马特恩核情形下达到最优阶。