Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50\,m average) and 3-second collision rate (0.18\%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.
翻译:视觉-语言-动作(VLA)模型能够描述场景并基于语言进行推理,但仍难以将其动作锚定于周围稠密的三维世界中。现有方法要么注入来自冻结三维基础模型的特征,却缺乏确保策略利用这些特征的训练目标;要么通过稀疏的边界框和地图损失约束几何结构,无法提供稠密的空间信号。我们提出VLGA,这是首个在监督下重建其行驶经过的稠密三维世界的视觉-语言-动作模型。VLGA通过专用专家模块将几何作为第四模态引入,与视觉、语言和动作并列,该专家通过逐像素点图回归损失(以激光雷达为基准)进行监督。在用于开环评估的nuScenes数据集和闭环评估的Bench2Drive数据集上开展的大量实验表明,VLGA相较于同类VLA方法的优越性。具体而言,在开环nuScenes数据集上,VLGA在无自车状态信息的VLA方法中创下新最优性能,达到最低L2误差(平均0.50米)和3秒碰撞率(0.18%)。在闭环Bench2Drive数据集上,VLGA以可比效率与舒适度取得79.08的最优驾驶得分,比此前最强VLA方法高出0.71。