Autonomous vehicle (AV) evaluation has been the subject of increased interest in recent years both in industry and in academia. This paper focuses on the development of a novel framework for generating adversarial driving behavior of background vehicle interfering against the AV to expose effective and rational risky events. Specifically, the adversarial behavior is learned by a reinforcement learning (RL) approach incorporated with the cumulative prospect theory (CPT) which allows representation of human risk cognition. Then, the extended version of deep deterministic policy gradient (DDPG) technique is proposed for training the adversarial policy while ensuring training stability as the CPT action-value function is leveraged. A comparative case study regarding the cut-in scenario is conducted on a high fidelity Hardware-in-the-Loop (HiL) platform and the results demonstrate the adversarial effectiveness to infer the weakness of the tested AV.
翻译:自动驾驶车辆评估近年来在工业界和学术界引起了广泛关注。本文聚焦于开发一种新框架,用于生成背景车辆干扰自动驾驶车辆的对抗性驾驶行为,以暴露有效且合理的危险事件。具体而言,对抗性行为通过结合累积前景理论的强化学习方法学习,该理论能够表征人类风险认知。随后,提出了扩展版本的深度确定性策略梯度技术用于训练对抗策略,同时利用累积前景理论动作价值函数确保训练稳定性。针对切入场景的对比案例研究在高保真硬件在环平台上进行,结果表明该方法能有效推断被测自动驾驶车辆的弱点。