2D forward-looking sonar is a crucial sensor for underwater robotic perception. A well-known problem in this field is estimating missing information in the elevation direction during sonar imaging. There are demands to estimate 3D information per image for 3D mapping and robot navigation during fly-through missions. Recent learning-based methods have demonstrated their strengths, but there are still drawbacks. Supervised learning methods have achieved high-quality results but may require further efforts to acquire 3D ground-truth labels. The existing self-supervised method requires pretraining using synthetic images with 3D supervision. This study aims to realize stable self-supervised learning of elevation angle estimation without pretraining using synthetic images. Failures during self-supervised learning may be caused by motion degeneracy problems. We first analyze the motion field of 2D forward-looking sonar, which is related to the main supervision signal. We utilize a modern learning framework and prove that if the training dataset is built with effective motions, the network can be trained in a self-supervised manner without the knowledge of synthetic data. Both simulation and real experiments validate the proposed method.
翻译:二维前视声纳是水下机器人感知的关键传感器。该领域的一个经典问题是在声纳成像过程中恢复仰角方向的缺失信息。在飞行穿越任务中,需要从每帧图像中估计三维信息以实现三维建图和机器人导航。近年来基于学习的方法已展现出优势,但仍存在局限。监督学习方法虽能获得高质量结果,但需要额外获取三维真实标签。现有的自监督方法需要利用带三维监督的合成图像进行预训练。本研究旨在实现无需合成图像预训练的自监督仰角估计稳定性学习。自监督学习过程中的失败可能源于运动退化问题。我们首先分析了与主要监督信号相关的二维前视声纳运动场,并利用现代学习框架证明:若训练数据集由有效运动构建,网络可在无合成数据的情况下进行自监督训练。仿真实验与真实数据实验均验证了所提方法的有效性。