Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can respond feasibly to autonomous vehicle (AV) behaviors that differ from the log. However, existing behavior simulation benchmarks do not directly measure reactive capability. They often let the simulator jointly control the AV and surrounding agents and evaluate realism through log similarity or open-loop prediction metrics. In this work, we introduce ReactSim-Bench for evaluating the reactive capability of behavior world model simulation in autonomous driving. We decouple the control of agents and the AV, using AV behaviors that differ from the log and require agents to respond as independent AV inputs. To obtain these AV behaviors, we construct a pipeline that uses an AV planner model to generate candidate behaviors and filters the data using rules and manual verification. Collision metrics, map-based metrics, and kinematic feasibility metrics are used to evaluate the safety and rule compliance of reactive responses. We construct 2,636 test scenarios with three categories and conduct a systematic evaluation of state-of-the-art models across multiple architectures, including Transformer-based, diffusion-based, and next-token-prediction-based models. We further analyze how replan frequency affects performance and provide insights for future studies.
翻译:反应能力是用于自动驾驶模拟系统的数据驱动行为世界模型模拟器的一个关键属性。具备这种能力后,模拟世界中的智能体能够对与日志记录不同的自动驾驶汽车行为做出可行的响应。然而,现有的行为模拟基准测试并未直接测量反应能力。它们通常让模拟器联合控制自动驾驶汽车和周围智能体,并通过日志相似性或开环预测指标来评估真实性。在这项工作中,我们引入了ReactSim-Bench,用于评估自动驾驶中行为世界模型模拟的反应能力。我们将智能体与自动驾驶汽车的控制解耦,使用与日志记录不同的自动驾驶汽车行为作为独立输入,要求智能体对此做出响应。为了获取这些自动驾驶汽车行为,我们构建了一个流水线,利用自动驾驶汽车规划器模型生成候选行为,并通过规则和手动验证对数据进行筛选。采用碰撞指标、基于地图的指标和运动学可行性指标来评估反应性响应的安全性和规则合规性。我们构建了包含三类共2636个测试场景,并对多个架构(包括基于Transformer、基于扩散和基于下一令牌预测的模型)中的最先进模型进行了系统评估。我们还进一步分析了重新规划频率如何影响性能,并为未来研究提供了见解。