Lidar depth completion is a new and hot topic of depth estimation. In this task, it is the key and difficult point to fuse the features of color space and depth space. In this paper, we migrate the classic LSTM and Transformer modules from NLP to depth completion and redesign them appropriately. Specifically, we use Forget gate, Update gate, Output gate, and Skip gate to achieve the efficient fusion of color and depth features and perform loop optimization at multiple scales. Finally, we further fuse the deep features through the Transformer multi-head attention mechanism. Experimental results show that without repetitive network structure and post-processing steps, our method can achieve state-of-the-art performance by adding our modules to a simple encoder-decoder network structure. Our method ranks first on the current mainstream autonomous driving KITTI benchmark dataset. It can also be regarded as a backbone network for other methods, which likewise achieves state-of-the-art performance.
翻译:激光雷达深度补全是深度估计领域一个新颖且热门的研究课题。在该任务中,融合颜色空间与深度空间的特征是核心难点。本文将经典的LSTM和Transformer模块从自然语言处理领域迁移至深度补全任务,并对其进行合理重构。具体而言,我们利用遗忘门、更新门、输出门和跳跃门实现颜色与深度特征的高效融合,并在多尺度上执行循环优化。最后,通过Transformer多头注意力机制进一步融合深层特征。实验结果表明,无需重复网络结构与后处理步骤,仅将所提模块嵌入简单的编码器-解码器架构即可达到最优性能。该方法在当前主流自动驾驶KITTI基准数据集上排名第一,同时可作为其他方法的骨干网络,同样取得最优性能。