Robots powered by 'blackbox' models need to provide human-understandable explanations which we can trust. Hence, explainability plays a critical role in trustworthy autonomous decision-making to foster transparency and acceptance among end users, especially in complex autonomous driving. Recent advancements in Multi-Modal Large Language models (MLLMs) have shown promising potential in enhancing the explainability as a driving agent by producing control predictions along with natural language explanations. However, severe data scarcity due to expensive annotation costs and significant domain gaps between different datasets makes the development of a robust and generalisable system an extremely challenging task. Moreover, the prohibitively expensive training requirements of MLLM and the unsolved problem of catastrophic forgetting further limit their generalisability post-deployment. To address these challenges, we present RAG-Driver, a novel retrieval-augmented multi-modal large language model that leverages in-context learning for high-performance, explainable, and generalisable autonomous driving. By grounding in retrieved expert demonstration, we empirically validate that RAG-Driver achieves state-of-the-art performance in producing driving action explanations, justifications, and control signal prediction. More importantly, it exhibits exceptional zero-shot generalisation capabilities to unseen environments without further training endeavours.
翻译:由“黑箱”模型驱动的机器人需要提供人类可理解的解释,以建立信任。因此,可解释性在可信自主决策中发挥着关键作用,有助于提高终端用户对复杂自动驾驶系统的透明度和接受度。多模态大语言模型的最新进展表明,这类模型通过生成控制预测及自然语言解释,有望增强作为驾驶主体的可解释性。然而,由于昂贵的标注成本导致数据严重匮乏,且不同数据集之间存在显著领域差异,开发鲁棒且通用的系统极具挑战性。此外,多模态大语言模型高昂的训练成本及灾难性遗忘这一未解难题,进一步限制了其部署后的泛化能力。针对上述挑战,我们提出RAG-Driver——一种基于检索增强的多模态大语言模型,通过利用上下文学习实现高性能、可解释且泛化的自动驾驶。实验证明,基于检索到的专家演示,RAG-Driver在生成驾驶行为解释、论证及控制信号预测方面达到了当前最优性能。更重要的是,它无需额外训练即可展现出对未见场景的卓越零样本泛化能力。