Imitation Learning (IL) has been increasingly employed to generate computationally efficient policies from task-relevant demonstrations provided by Model Predictive Control (MPC). However, commonly employed IL methods are often data- and computationally-inefficient, as they require a large number of MPC demonstrations, resulting in long training times, and they produce policies with limited robustness to disturbances not experienced during training. In this work, we propose an IL strategy to efficiently compress a computationally expensive MPC into a Deep Neural Network (DNN) policy that is robust to previously unseen disturbances. By using a robust variant of the MPC, called Robust Tube MPC (RTMPC), and leveraging properties from the controller, we introduce a computationally-efficient Data Aggregation (DA) method that enables a significant reduction of the number of MPC demonstrations and training time required to generate a robust policy. Our approach opens the possibility of zero-shot transfer of a policy trained from a single MPC demonstration collected in a nominal domain, such as a simulation or a robot in a lab/controlled environment, to a new domain with previously-unseen bounded model errors/perturbations. Numerical and experimental evaluations performed using linear and nonlinear MPC for agile flight on a multirotor show that our method outperforms strategies commonly employed in IL (such as DAgger and DR) in terms of demonstration-efficiency, training time, and robustness to perturbations unseen during training.
翻译:模仿学习(IL)越来越多地被用于从模型预测控制(MPC)提供的任务相关示范中生成计算高效的策略。然而,常用的IL方法通常数据效率低且计算效率低,因为它们需要大量MPC示范,导致训练时间长,并且生成的策略对训练中未经历的扰动鲁棒性有限。在这项工作中,我们提出了一种IL策略,用于将计算昂贵的MPC高效压缩为深度神经网络(DNN)策略,该策略对先前未见过的扰动具有鲁棒性。通过使用MPC的鲁棒变体——鲁棒管MPC(RTMPC)并利用控制器的特性,我们引入了一种计算高效的数据聚合(DA)方法,该方法能够显著减少生成鲁棒策略所需的MPC示范数量和训练时间。我们的方法开辟了将基于名义域(如仿真或实验室/受控环境中的机器人)收集的单个MPC示范训练的策略零样本迁移到具有先前未见过的有界模型误差/扰动的新域的可能性。使用线性与非线性MPC在多旋翼快速飞行中进行的数值与实验评估表明,我们的方法在示范效率、训练时间以及对训练中未见扰动的鲁棒性方面优于IL中常用的策略(如DAgger和DR)。