The deployment of agile autonomous systems in challenging, unstructured environments requires adaptation capabilities and robustness to uncertainties. Existing robust and adaptive controllers, such as the ones based on MPC, can achieve impressive performance at the cost of heavy online onboard computations. Strategies that efficiently learn robust and onboard-deployable policies from MPC have emerged, but they still lack fundamental adaptation capabilities. In this work, we extend an existing efficient IL algorithm for robust policy learning from MPC with the ability to learn policies that adapt to challenging model/environment uncertainties. The key idea of our approach consists in modifying the IL procedure by conditioning the policy on a learned lower-dimensional model/environment representation that can be efficiently estimated online. We tailor our approach to the task of learning an adaptive position and attitude control policy to track trajectories under challenging disturbances on a multirotor. Our evaluation is performed in a high-fidelity simulation environment and shows that a high-quality adaptive policy can be obtained in about $1.3$ hours. We additionally empirically demonstrate rapid adaptation to in- and out-of-training-distribution uncertainties, achieving a $6.1$ cm average position error under a wind disturbance that corresponds to about $50\%$ of the weight of the robot and that is $36\%$ larger than the maximum wind seen during training.
翻译:在挑战性、非结构化环境中部署敏捷自主系统需要适应能力和对不确定性的鲁棒性。现有的鲁棒自适应控制器(例如基于MPC的控制器)可以以沉重的在线机载计算为代价实现令人印象深刻的性能。从MPC高效学习鲁棒且可机载部署策略的方法已经出现,但仍缺乏基本的适应能力。在这项工作中,我们将一种现有的从MPC进行鲁棒策略学习的高效IL算法进行扩展,使其能够学习适应具有挑战性的模型/环境不确定性的策略。我们方法的关键思想在于通过将策略条件化为一个低维模型/环境表示来修改IL过程,该表示可在线上高效估计。我们将该方法应用于在四旋翼飞行器上学习自适应位置和姿态控制策略,以在具有挑战性的扰动下跟踪轨迹。在高保真仿真环境中进行的评估表明,可在约1.3小时内获得高质量的自适应策略。此外,我们通过实验证明了该策略能快速适应训练分布内和训练分布外的不确定性,在相当于机器人重量约50%且比训练时观测到的最大风速大36%的风扰动下,实现了6.1厘米的平均位置误差。