Sampling-based Model Predictive Control (MPC) has been a practical and effective approach in many domains, notably model-based reinforcement learning, thanks to its flexibility and parallelizability. Despite its appealing empirical performance, the theoretical understanding, particularly in terms of convergence analysis and hyperparameter tuning, remains absent. In this paper, we characterize the convergence property of a widely used sampling-based MPC method, Model Predictive Path Integral Control (MPPI). We show that MPPI enjoys at least linear convergence rates when the optimization is quadratic, which covers time-varying LQR systems. We then extend to more general nonlinear systems. Our theoretical analysis directly leads to a novel sampling-based MPC algorithm, CoVariance-Optimal MPC (CoVo-MPC) that optimally schedules the sampling covariance to optimize the convergence rate. Empirically, CoVo-MPC significantly outperforms standard MPPI by 43-54% in both simulations and real-world quadrotor agile control tasks. Videos and Appendices are available at \url{https://lecar-lab.github.io/CoVO-MPC/}.
翻译:基于采样的模型预测控制(MPC)因其灵活性和可并行化特性,已成为许多领域(尤其是基于模型的强化学习)中实用且有效的方法。尽管其经验性能令人瞩目,但其理论理解,特别是在收敛性分析和超参数调优方面,仍然缺失。本文刻画了广泛使用的基于采样的MPC方法——模型预测路径积分控制(MPPI)的收敛性质。我们证明,当优化问题是二次型时(涵盖时变LQR系统),MPPI至少具有线性收敛速率。随后,我们将分析推广至更一般的非线性系统。我们的理论分析直接催生了一种新颖的基于采样的MPC算法——协方差最优MPC(CoVO-MPC),该算法通过最优调度采样协方差来优化收敛速率。实验表明,在仿真和真实世界四旋翼敏捷控制任务中,CoVO-MPC相比标准MPPI显著提升43-54%。视频与附录详见\url{https://lecar-lab.github.io/CoVO-MPC/}。