We investigate the problem of autonomous racing among teams of cooperative agents that are subject to realistic racing rules. Our work extends previous research on hierarchical control in head-to-head autonomous racing by considering a generalized version of the problem while maintaining the two-level hierarchical control structure. A high-level tactical planner constructs a discrete game that encodes the complex rules using simplified dynamics to produce a sequence of target waypoints. The low-level path planner uses these waypoints as a reference trajectory and computes high-resolution control inputs by solving a simplified formulation of a racing game with a simplified representation of the realistic racing rules. We explore two approaches for the low-level path planner: training a multi-agent reinforcement learning (MARL) policy and solving a linear-quadratic Nash game (LQNG) approximation. We evaluate our controllers on simple and complex tracks against three baselines: an end-to-end MARL controller, a MARL controller tracking a fixed racing line, and an LQNG controller tracking a fixed racing line. Quantitative results show our hierarchical methods outperform the baselines in terms of race wins, overall team performance, and compliance with the rules. Qualitatively, we observe the hierarchical controllers mimic actions performed by expert human drivers such as coordinated overtaking, defending against multiple opponents, and long-term planning for delayed advantages.
翻译:我们研究了受限于真实竞技规则的协作智能体团队在自主竞速场景下的问题。在保持双层分层控制结构的同时,通过考虑该问题的广义形式,我们将此前头对头自主竞速中分层控制的研究进行了扩展。高层战术规划器利用简化动力学构建编码复杂规则的离散博弈,生成一系列目标路径点。低层路径规划器将这些路径点作为参考轨迹,通过求解带有简化真实竞赛规则的竞速博弈简化公式,计算高分辨率控制输入。我们探索了低层路径规划器的两种实现方法:训练多智能体强化学习(MARL)策略,以及求解线性二次纳什博弈(LQNG)近似。我们在简单与复杂赛道上,将所提控制器与三种基线方法进行对比:端到端MARL控制器、追踪固定赛线的MARL控制器,以及追踪固定赛线的LQNG控制器。定量结果表明,在获胜次数、团队整体表现及规则遵守度方面,我们的分层方法均优于基线方法。定性分析显示,分层控制器能够模拟人类专家驾驶员的操作行为,例如协同超车、防御多辆对手车辆,以及为延迟性优势进行长期规划。