Autonomous Racing has seen remarkable progress through deep Reinforcement Learning (RL), primarily for four-wheeled vehicles. However, motorbikes introduce substantially greater complexity due to the need to manage balance and lean angle, in addition to more reactive steering and throttle control, and a smaller weight. In this work, we present a framework for training an autonomous agent to race a superbike in VRider SBK, a physics-accurate Unity-based motorbike simulator. Our approach integrates Soft Actor-Critic (SAC) with Self-Paced curriculum Deep reinforcement Learning (SPDL), which dynamically generates progressively more challenging tasks based on the agent's performance, without requiring manual curriculum design. The agent's state space comprises proprioceptive features extended with lean-angle history, along with global track features via course points. The reward signal is shaped to encourage progress along the track while penalizing instability-inducing behaviors specific to two-wheeled dynamics. Preliminary experimental results demonstrate that SPDL outperforms SAC alone in training efficiency, lap time, and driving stability across multiple tracks and motorbike models, establishing a first baseline for RL-based autonomous motorbike racing.
翻译:自动驾驶赛车通过深度强化学习取得了显著进展,主要针对四轮车辆。然而,摩托车由于需要管理平衡和倾斜角度,再加上更灵敏的转向和油门控制以及更小的重量,带来了更高的复杂性。在这项工作中,我们提出一个框架,用于训练自主智能体在基于Unity的物理精确摩托车模拟器VRider SBK中比赛超级摩托车。我们的方法将软演员-评论家与自定进度深度强化学习课程相结合,该课程根据智能体的表现动态生成逐渐更具挑战性的任务,无需手动设计课程。智能体的状态空间包括扩展了倾斜角度历史的本体感知特征,以及通过赛道点提供的全局赛道特征。奖励信号设计为鼓励沿赛道前进,同时惩罚特定于两轮动力学的不稳定行为。初步实验结果表明,SPDL在多个赛道和摩托车模型上的训练效率、圈速和驾驶稳定性方面均优于单独的SAC,为基于强化学习的自主摩托车比赛建立了第一个基线。