Model predictive control (MPC) has been applied to many platforms in robotics and autonomous systems for its capability to predict a system's future behavior while incorporating constraints that a system may have. To enhance the performance of a system with an MPC controller, one can manually tune the MPC's cost function. However, it can be challenging due to the possibly high dimension of the parameter space as well as the potential difference between the open-loop cost function in MPC and the overall closed-loop performance metric function. This paper presents DiffTune-MPC, a novel learning method, to learn the cost function of an MPC in a closed-loop manner. The proposed framework is compatible with the scenario where the time interval for performance evaluation and MPC's planning horizon have different lengths. We show the auxiliary problem whose solution admits the analytical gradients of MPC and discuss its variations in different MPC settings, including nonlinear MPCs that are solved using sequential quadratic programming. Simulation results demonstrate the learning capability of DiffTune-MPC and the generalization capability of the learned MPC parameters.
翻译:模型预测控制(MPC)因其能够预测系统未来行为并同时考虑系统可能存在的约束,已被广泛应用于机器人学与自主系统中的多种平台。为提升采用MPC控制器的系统性能,通常需手动调节MPC的成本函数。然而,由于参数空间可能维度较高,且MPC中的开环成本函数与整体闭环性能评价函数之间可能存在差异,这一调节过程往往具有挑战性。本文提出DiffTune-MPC这一新颖的学习方法,以闭环方式学习MPC的成本函数。所提框架兼容性能评价时间区间与MPC规划视野长度不同的场景。我们给出了其解可导出MPC解析梯度的辅助问题,并讨论了该方法在不同MPC设置下的变体,包括采用序列二次规划求解的非线性MPC。仿真结果验证了DiffTune-MPC的学习能力以及所学MPC参数的泛化能力。