Head-to-head autonomous racing is a challenging problem, as the vehicle needs to operate at the friction or handling limits in order to achieve minimum lap times while also actively looking for strategies to overtake/stay ahead of the opponent. In this work we propose a head-to-head racing environment for reinforcement learning which accurately models vehicle dynamics. Some previous works have tried learning a policy directly in the complex vehicle dynamics environment but have failed to learn an optimal policy. In this work, we propose a curriculum learning-based framework by transitioning from a simpler vehicle model to a more complex real environment to teach the reinforcement learning agent a policy closer to the optimal policy. We also propose a control barrier function-based safe reinforcement learning algorithm to enforce the safety of the agent in a more effective way while not compromising on optimality.
翻译:头对头无人赛车是一个具有挑战性的问题,因为车辆需要在摩擦或操控极限状态下运行,以实现最短圈时,同时积极寻找超车或保持领先对手的策略。本文提出一种用于强化学习的头对头赛车环境,该环境精确建模了车辆动力学。以往一些工作尝试直接在复杂车辆动力学环境中学习策略,但未能学到最优策略。本研究提出一种基于课程学习的框架,通过从简单车辆模型过渡到更复杂的真实环境,使强化学习智能体学得更接近最优策略的策略。我们还提出一种基于控制障碍函数的安全强化学习算法,在更有效保障智能体安全性的同时不牺牲最优性。