We revisit in this paper the discrete-time linear quadratic regulator (LQR) problem from the perspective of receding-horizon policy gradient (RHPG), a newly developed model-free learning framework for control applications. We provide a fine-grained sample complexity analysis for RHPG to learn a control policy that is both stabilizing and $\epsilon$-close to the optimal LQR solution, and our algorithm does not require knowing a stabilizing control policy for initialization. Combined with the recent application of RHPG in learning the Kalman filter, we demonstrate the general applicability of RHPG in linear control and estimation with streamlined analyses.
翻译:本文从滚动时域策略梯度(RHPG)这一新近发展的无模型控制学习框架出发,重新审视了离散时间线性二次型调节器(LQR)问题。我们针对RHPG算法给出了精细的样本复杂度分析,使其能够学习到既稳定又与原最优LQR解相差不超过ϵ的控制策略,且该算法无需预先已知稳定的初始化控制策略。结合RHPG近期在卡尔曼滤波器学习中的应用,我们通过简化分析流程,展示了RHPG在线性控制与估计问题中的广泛适用性。