We present a system that enables an autonomous small-scale RC car to drive aggressively from visual observations using reinforcement learning (RL). Our system, FastRLAP (faster lap), trains autonomously in the real world, without human interventions, and without requiring any simulation or expert demonstrations. Our system integrates a number of important components to make this possible: we initialize the representations for the RL policy and value function from a large prior dataset of other robots navigating in other environments (at low speed), which provides a navigation-relevant representation. From here, a sample-efficient online RL method uses a single low-speed user-provided demonstration to determine the desired driving course, extracts a set of navigational checkpoints, and autonomously practices driving through these checkpoints, resetting automatically on collision or failure. Perhaps surprisingly, we find that with appropriate initialization and choice of algorithm, our system can learn to drive over a variety of racing courses with less than 20 minutes of online training. The resulting policies exhibit emergent aggressive driving skills, such as timing braking and acceleration around turns and avoiding areas which impede the robot's motion, approaching the performance of a human driver using a similar first-person interface over the course of training.
翻译:我们提出了一套系统,使自主小型遥控赛车能够通过强化学习从视觉观测中实现激进驾驶。我们的系统FastRLAP(更快圈速)在真实世界中自主训练,无需人工干预,也不依赖任何仿真或专家示范。为实现这一目标,系统集成了多项关键组件:从其他机器人在不同环境中(低速)导航的大规模先验数据集中,初始化强化学习策略与价值函数的表征,从而获得具备导航相关性的特征表达。在此基础上,一种样本高效的在线强化学习方法利用单次低速用户提供的示范来确定所需驾驶路线,提取一组导航路径点,并自主练习穿越这些路径点,在碰撞或失败时自动重置。令人惊讶的是,我们发现通过适当的初始化与算法选择,该系统可在不到20分钟的在线训练中学会在多种赛道上驾驶。最终生成的策略展现出涌现性的激进驾驶技能,例如在弯道处精准控制刹车与加速时机、避开阻碍机器人运动的区域,其性能在训练过程中已接近使用类似第一人称界面的人类驾驶员水平。