Reinforcement learning (RL) has shown empirical success in various real world settings with complex models and large state-action spaces. The existing analytical results, however, typically focus on settings with a small number of state-actions or simple models such as linearly modeled state-action value functions. To derive RL policies that efficiently handle large state-action spaces with more general value functions, some recent works have considered nonlinear function approximation using kernel ridge regression. We propose $\pi$-KRVI, an optimistic modification of least-squares value iteration, when the state-action value function is represented by a reproducing kernel Hilbert space (RKHS). We prove the first order-optimal regret guarantees under a general setting. Our results show a significant polynomial in the number of episodes improvement over the state of the art. In particular, with highly non-smooth kernels (such as Neural Tangent kernel or some Mat\'ern kernels) the existing results lead to trivial (superlinear in the number of episodes) regret bounds. We show a sublinear regret bound that is order optimal in the case of Mat\'ern kernels where a lower bound on regret is known.
翻译:强化学习在复杂模型和大状态-动作空间的各类实际场景中已展现出经验成功,但现有解析结果通常聚焦于少量状态-动作或线性建模状态-动作值函数等简单设定。为推导能高效处理大状态-动作空间并采用更一般值函数的强化学习策略,近期工作利用核岭回归实现了非线性函数逼近。我们提出π-KRVI——一种基于最小二乘值迭代的乐观改进算法,其作用于再生核希尔伯特空间(RKHS)表示的状态-动作值函数。我们首次在一般设定下证明了阶最优遗憾界保证,该结果显示在回合数维度上较现有最先进方法具有显著多项式改进。特别地,对于高度非光滑核函数(如神经正切核或某些Matérn核),现有方法仅能获得平凡(超线性于回合数)的遗憾界,而我们所证明的次线性遗憾界在已知下界的Matérn核情形下达到阶最优。