As a fundamental task of vision-based perception, 3D occupancy prediction reconstructs 3D structures of surrounding environments. It provides detailed information for autonomous driving planning and navigation. However, most existing methods heavily rely on the LiDAR point clouds to generate occupancy ground truth, which is not available in the vision-based system. In this paper, we propose an OccNeRF method for self-supervised multi-camera occupancy prediction. Different from bounded 3D occupancy labels, we need to consider unbounded scenes with raw image supervision. To solve the issue, we parameterize the reconstructed occupancy fields and reorganize the sampling strategy. The neural rendering is adopted to convert occupancy fields to multi-camera depth maps, supervised by multi-frame photometric consistency. Moreover, for semantic occupancy prediction, we design several strategies to polish the prompts and filter the outputs of a pretrained open-vocabulary 2D segmentation model. Extensive experiments for both self-supervised depth estimation and semantic occupancy prediction tasks on nuScenes dataset demonstrate the effectiveness of our method.
翻译:作为基于视觉感知的基础任务,三维占用预测能够重建周围环境的三维结构,为自动驾驶规划与导航提供详细信息。然而,现有方法大多严重依赖激光雷达点云生成占用真值,这在纯视觉系统中无法获取。本文提出一种名为OccNeRF的自监督多视角占用预测方法。与有界的三维占用标签不同,我们需要考虑通过原始图像监督的无界场景。为解决这一问题,我们对重建的占用场进行参数化并重新组织采样策略。采用神经渲染技术将占用场转换为多视角深度图,并通过多帧光度一致性进行监督。此外,针对语义占用预测任务,我们设计了多种策略来优化提示词并过滤预训练开放词汇二维分割模型的输出。在nuScenes数据集上的自监督深度估计与语义占用预测任务的广泛实验证明了我们方法的有效性。