This paper studies convergence rates for some value function approximations that arise in a collection of reproducing kernel Hilbert spaces (RKHS) $H(\Omega)$. By casting an optimal control problem in a specific class of native spaces, strong rates of convergence are derived for the operator equation that enables offline approximations that appear in policy iteration. Explicit upper bounds on error in value function approximations are derived in terms of power function $\Pwr_{H,N}$ for the space of finite dimensional approximants $H_N$ in the native space $H(\Omega)$. These bounds are geometric in nature and refine some well-known, now classical results concerning convergence of approximations of value functions.
翻译:本文研究了一类再生核希尔伯特空间(RKHS)$H(\Omega)$中出现的若干价值函数近似的收敛速率。通过将最优控制问题置于特定类别的本征空间中,推导了操作方程在策略迭代中实现离线近似时的强收敛速率。价值函数近似误差的显式上界由有限维逼近空间$H_N$在本征空间$H(\Omega)$中的幂函数$\Pwr_{H,N}$表达。这些几何性质的界改进了部分关于价值函数近似收敛性的经典结论。