Large vision-language models (VLMs) have garnered increasing interest in autonomous driving areas, due to their advanced capabilities in complex reasoning tasks essential for highly autonomous vehicle behavior. Despite their potential, research in autonomous systems is hindered by the lack of datasets with annotated reasoning chains that explain the decision-making processes in driving. To bridge this gap, we present Reason2Drive, a benchmark dataset with over 600K video-text pairs, aimed at facilitating the study of interpretable reasoning in complex driving environments. We distinctly characterize the autonomous driving process as a sequential combination of perception, prediction, and reasoning steps, and the question-answer pairs are automatically collected from a diverse range of open-source outdoor driving datasets, including nuScenes, Waymo and ONCE. Moreover, we introduce a novel aggregated evaluation metric to assess chain-based reasoning performance in autonomous systems, addressing the semantic ambiguities of existing metrics such as BLEU and CIDEr. Based on the proposed benchmark, we conduct experiments to assess various existing VLMs, revealing insights into their reasoning capabilities. Additionally, we develop an efficient approach to empower VLMs to leverage object-level perceptual elements in both feature extraction and prediction, further enhancing their reasoning accuracy. The code and dataset will be released.
翻译:大型视觉语言模型(VLMs)因其在高度自动化驾驶行为所需的复杂推理任务中展现出的卓越能力,在自动驾驶领域日益受到关注。尽管潜力巨大,但由于缺乏标注驾驶决策过程推理链的数据集,自主系统的研究仍面临阻碍。为弥补这一空白,我们提出Reason2Drive基准数据集,包含超过60万条视频-文本对,旨在促进复杂驾驶环境中可解释推理的研究。我们将自动驾驶过程明确表征为感知、预测和推理步骤的序列组合,并从nuScenes、Waymo和ONCE等多样化的开源户外驾驶数据集中自动收集问答对。此外,我们引入一种新颖的聚合评估指标,用于评估自主系统中的链式推理性能,以解决BLEU和CIDEr等现有指标存在的语义歧义问题。基于所提出的基准,我们开展实验评估多种现有VLM,揭示其推理能力的洞见。同时,我们开发了一种高效方法,使VLM能够在特征提取和预测中利用对象级感知元素,进一步提升推理准确性。相关代码与数据集将公开提供。