This work evaluates and analyzes the combination of imitation learning (IL) and differentiable model predictive control (MPC) for the application of human-like autonomous driving. We combine MPC with a hierarchical learning-based policy, and measure its performance in open-loop and closed-loop with metrics related to safety, comfort and similarity to human driving characteristics. We also demonstrate the value of augmenting open-loop behavioral cloning with closed-loop training for a more robust learning, approximating the policy gradient through time with the state space model used by the MPC. We perform experimental evaluations on a lane keeping control system, learned from demonstrations collected on a fixed-base driving simulator, and show that our imitative policies approach the human driving style preferences.
翻译:本文评估并分析了将模仿学习(IL)与可微分模型预测控制(MPC)相结合用于类人自动驾驶的应用。我们将MPC与分层学习策略相结合,并从安全性、舒适性和与人类驾驶特征的相似性等指标出发,在开环和闭环条件下测量其性能。我们还展示了利用闭环训练增强开环行为克隆的价值,以增强学习的鲁棒性,通过MPC使用的状态空间模型近似随时间变化的策略梯度。我们在车道保持控制系统上进行了实验评估,该系统基于固定基座驾驶模拟器采集的示范数据学习,结果表明我们的模仿策略能够接近人类驾驶风格偏好。