While autonomous vehicles (AVs) may perform remarkably well in generic real-life cases, their irrational action in some unforeseen cases leads to critical safety concerns. This paper introduces the concept of collaborative reinforcement learning (RL) to generate challenging test cases for AV planning and decision-making module. One of the critical challenges for collaborative RL is the credit assignment problem, where a proper assignment of rewards to multiple agents interacting in the traffic scenario, considering all parameters and timing, turns out to be non-trivial. In order to address this challenge, we propose a novel potential-based reward-shaping approach inspired by counterfactual analysis for solving the credit-assignment problem. The evaluation in a simulated environment demonstrates the superiority of our proposed approach against other methods using local and global rewards.
翻译:尽管自动驾驶车辆(AV)在常规真实场景中表现优异,但在某些不可预见情况下其非理性行为仍会引发严重安全隐患。本文提出协作强化学习(RL)概念,用于生成具有挑战性的测试用例以检验自动驾驶规划与决策模块。协作强化学习面临的关键挑战之一是信用分配问题——需综合考虑交通场景中所有参数与时序因素,对多智能体交互进行合理奖励分配,这显然极具难度。为应对这一挑战,我们提出一种受反事实分析启发的基于势能的奖励塑形方法,用于解决信用分配问题。仿真环境中的评估结果表明,相较于使用局部与全局奖励的其他方法,本方法具有显著优越性。