Driving in dense traffic with human and autonomous drivers is a challenging task that requires high level planning and reasoning along with the ability to react quickly to changes in a dynamic environment. In this study, we propose a hierarchical learning approach that uses learned motion primitives as actions. Motion primitives are obtained using unsupervised skill discovery without a predetermined reward function, allowing them to be reused in different scenarios. This can reduce the total training time for applications that need to obtain multiple models with varying behavior. Simulation results demonstrate that the proposed approach yields driver models that achieve higher performance with less training compared to baseline reinforcement learning methods.
翻译:在存在人类驾驶员和自主驾驶员的密集交通环境中进行驾驶是一项极具挑战性的任务,它既需要高层次的规划与推理能力,又需要能够对动态环境中的变化做出快速反应。本研究提出了一种分层学习方法,该方法将学习到的运动基元作为动作。这些运动基元通过无监督技能发现技术获得,无需预设奖励函数,从而使得它们能够在不同场景中被重复利用。这一特性可以减少需要获得多种不同行为模型的应用程序的总训练时间。仿真结果表明,与基准强化学习方法相比,所提出的方法能够以更少的训练量获得性能更高的驾驶员模型。