We propose a novel knowledge distillation framework for effectively teaching a sensorimotor student agent to drive from the supervision of a privileged teacher agent. Current distillation for sensorimotor agents methods tend to result in suboptimal learned driving behavior by the student, which we hypothesize is due to inherent differences between the input, modeling capacity, and optimization processes of the two agents. We develop a novel distillation scheme that can address these limitations and close the gap between the sensorimotor agent and its privileged teacher. Our key insight is to design a student which learns to align their input features with the teacher's privileged Bird's Eye View (BEV) space. The student then can benefit from direct supervision by the teacher over the internal representation learning. To scaffold the difficult sensorimotor learning task, the student model is optimized via a student-paced coaching mechanism with various auxiliary supervision. We further propose a high-capacity imitation learned privileged agent that surpasses prior privileged agents in CARLA and ensures the student learns safe driving behavior. Our proposed sensorimotor agent results in a robust image-based behavior cloning agent in CARLA, improving over current models by over 20.6% in driving score without requiring LiDAR, historical observations, ensemble of models, on-policy data aggregation or reinforcement learning.
翻译:我们提出了一种新颖的知识蒸馏框架,旨在通过特权教师智能体的监督,有效训练感知运动学生智能体执行驾驶任务。当前面向感知运动智能体的蒸馏方法往往导致学生学到的驾驶行为次优,我们假设这是由于两个智能体在输入、建模能力及优化过程中的固有差异所致。我们开发了一种新颖的蒸馏方案,能够解决这些局限性,缩小感知运动智能体与其特权教师之间的差距。关键洞察在于设计一个学生模型,使其学习将输入特征与教师的特权鸟瞰图空间对齐。这样,学生便能在内部表征学习过程中获得教师的直接监督。为支撑这一困难的感知运动学习任务,学生模型通过一种学生驱动的渐进式指导机制进行优化,结合多种辅助监督。我们进一步提出了一种高容量模仿学习特权智能体,该智能体在CARLA仿真环境中超越了先前的特权智能体,并确保学生学得安全的驾驶行为。我们提出的感知运动智能体在CARLA中构建了鲁棒的基于图像的模仿学习智能体,在不依赖激光雷达、历史观测、模型集成、在线策略数据聚合或强化学习的情况下,驾驶评分比当前模型提升了超过20.6%。