Vision-based autonomous driving has gained much attention due to its low costs and excellent performance. Compared with dense BEV (Bird's Eye View) or sparse query models, Gaussian-centric method is a comprehensive yet sparse representation by describing scene with 3D semantic Gaussians. In this paper, we introduce DLWM, a novel paradigm with Dual Latent World Models specifically designed to enable holistic gaussian-centric pre-training in autonomous driving using two stages. In the first stage, DLWM predicts 3D Gaussians from queries by self-supervised reconstructing multi-view semantic and depth images. Equipped with fine-grained contextual features, in the second stage, two latent world models are trained separately for temporal feature learning, including Gaussian-flow-guided latent prediction for downstream occupancy perception and forecasting tasks, and ego-planning-guided latent prediction for motion planning. Extensive experiments in SurroundOcc and nuScenes benchmarks demonstrate that DLWM shows significant performance gains across Gaussian-centric 3D occupancy perception, 4D occupancy forecasting and motion planning tasks.
翻译:摘要:基于视觉的自动驾驶因其低成本与优异性能而备受关注。相较于稠密的BEV(鸟瞰图)或稀疏查询模型,高斯中心方法通过以三维语义高斯描述场景,成为一种综合且稀疏的表示方式。本文提出DLWM——一种新型双潜世界模型范式,专为自动驾驶中的全局高斯中心预训练设计,采用两阶段架构。第一阶段中,DLWM通过自监督重构多视图语义与深度图像,从查询中预测三维高斯。配备细粒度上下文特征后,第二阶段分别训练两个潜世界模型以进行时序特征学习:其一是面向下游占据感知与预测任务的高斯流引导潜变量预测,其二是面向运动规划的自我规划引导潜变量预测。在SurroundOcc与nuScenes基准上的大量实验表明,DLWM在三维高斯中心占据感知、四维占据预测及运动规划任务中均展现出显著性能提升。