An open problem in autonomous vehicle safety validation is building reliable models of human driving behavior in simulation. This work presents an approach to learn neural driving policies from real world driving demonstration data. We model human driving as a sequential decision making problem that is characterized by non-linearity and stochasticity, and unknown underlying cost functions. Imitation learning is an approach for generating intelligent behavior when the cost function is unknown or difficult to specify. Building upon work in inverse reinforcement learning (IRL), Generative Adversarial Imitation Learning (GAIL) aims to provide effective imitation even for problems with large or continuous state and action spaces, such as modeling human driving. This article describes the use of GAIL for learning-based driver modeling. Because driver modeling is inherently a multi-agent problem, where the interaction between agents needs to be modeled, this paper describes a parameter-sharing extension of GAIL called PS-GAIL to tackle multi-agent driver modeling. In addition, GAIL is domain agnostic, making it difficult to encode specific knowledge relevant to driving in the learning process. This paper describes Reward Augmented Imitation Learning (RAIL), which modifies the reward signal to provide domain-specific knowledge to the agent. Finally, human demonstrations are dependent upon latent factors that may not be captured by GAIL. This paper describes Burn-InfoGAIL, which allows for disentanglement of latent variability in demonstrations. Imitation learning experiments are performed using NGSIM, a real-world highway driving dataset. Experiments show that these modifications to GAIL can successfully model highway driving behavior, accurately replicating human demonstrations and generating realistic, emergent behavior in the traffic flow arising from the interaction between driving agents.
翻译:自动驾驶安全性验证中的一个开放问题是在仿真中构建可靠的人类驾驶行为模型。本文提出了一种从真实驾驶演示数据中学习神经驾驶策略的方法。我们将人类驾驶建模为一个以非线性、随机性和未知底层代价函数为特征的序列决策问题。模仿学习是一种在代价函数未知或难以指定时生成智能行为的方法。基于逆强化学习(IRL)的研究,生成对抗模仿学习(GAIL)旨在为具有大规模或连续状态与动作空间的问题(如人类驾驶建模)提供有效模仿。本文描述了使用GAIL进行基于学习的驾驶员建模。由于驾驶员建模本质上是一个多智能体问题,需要建模智能体之间的交互,本文提出了一种GAIL的参数共享扩展——PS-GAIL,以解决多智能体驾驶员建模问题。此外,GAIL是领域无关的,难以在学习过程中编码与驾驶相关的特定知识。本文描述了奖励增强模仿学习(RAIL),该方法通过修改奖励信号向智能体提供领域特定知识。最后,人类演示依赖于GAIL可能无法捕捉到的潜在因素。本文介绍了Burn-InfoGAIL,它能够解耦演示中的潜在变量。模仿学习实验使用真实高速公路驾驶数据集NGSIM进行。实验表明,这些对GAIL的改进能够成功建模高速公路驾驶行为,准确复现人类演示,并在驾驶智能体交互产生的交通流中生成逼真的涌现行为。