The recent advances in camera-based bird's eye view (BEV) representation exhibit great potential for in-vehicle 3D perception. Despite the substantial progress achieved on standard benchmarks, the robustness of BEV algorithms has not been thoroughly examined, which is critical for safe operations. To bridge this gap, we introduce RoboBEV, a comprehensive benchmark suite that encompasses eight distinct corruptions, including Bright, Dark, Fog, Snow, Motion Blur, Color Quant, Camera Crash, and Frame Lost. Based on it, we undertake extensive evaluations across a wide range of BEV-based models to understand their resilience and reliability. Our findings indicate a strong correlation between absolute performance on in-distribution and out-of-distribution datasets. Nonetheless, there are considerable variations in relative performance across different approaches. Our experiments further demonstrate that pre-training and depth-free BEV transformation has the potential to enhance out-of-distribution robustness. Additionally, utilizing long and rich temporal information largely helps with robustness. Our findings provide valuable insights for designing future BEV models that can achieve both accuracy and robustness in real-world deployments.
翻译:摘要:近年来,基于摄像头的鸟瞰图(BEV)表示方法在车载3D感知领域展现出巨大潜力。尽管在标准基准上取得了显著进展,但BEV算法的鲁棒性尚未得到充分检验,而这对于安全运行至关重要。为弥补这一空白,我们提出了RoboBEV,一个全面的基准测试套件,涵盖了八种不同的损坏类型,包括亮度、黑暗、雾、雪、运动模糊、颜色量化、摄像头崩溃和帧丢失。在此基础上,我们对多种基于BEV的模型进行了广泛评估,以了解其韧性和可靠性。我们的研究结果表明,在分布内和分布外数据集上的绝对性能之间存在强相关性,然而,不同方法之间的相对性能存在显著差异。我们的实验进一步证明,预训练和无深度的BEV变换技术有潜力增强分布外的鲁棒性。此外,利用丰富且长时间的时间信息在很大程度上有助于提升鲁棒性。我们的发现为设计未来能在实际部署中同时实现准确性和鲁棒性的BEV模型提供了宝贵见解。