Visual place recognition (VPR) enables autonomous systems to localize themselves within an environment using image information. While Convolution Neural Networks (CNNs) currently dominate state-of-the-art VPR performance, their high computational requirements make them unsuitable for platforms with budget or size constraints. This has spurred the development of lightweight algorithms, such as DrosoNet, which employs a voting system based on multiple bio-inspired units. In this paper, we present a novel training approach for DrosoNet, wherein separate models are trained on distinct regions of a reference image, allowing them to specialize in the visual features of that specific section. Additionally, we introduce a convolutional-like prediction method, in which each DrosoNet unit generates a set of place predictions for each portion of the query image. These predictions are then combined using the previously introduced voting system. Our approach significantly improves upon the VPR performance of previous work while maintaining an extremely compact and lightweight algorithm, making it suitable for resource-constrained platforms.
翻译:摘要:视觉地点识别(VPR)使自主系统能够利用图像信息在环境中进行定位。尽管卷积神经网络(CNN)目前主导着最先进的VPR性能,但其高计算需求使其不适用于预算或尺寸受限的平台。这推动了轻量级算法的发展,例如DrosoNet,它采用基于多个仿生单元的投票系统。在本文中,我们提出了一种针对DrosoNet的新型训练方法,其中针对参考图像的不同区域分别训练独立模型,使其能够专注于该特定区域的视觉特征。此外,我们引入了一种类卷积预测方法,每个DrosoNet单元针对查询图像的每个部分生成一组地点预测。这些预测随后通过先前提出的投票系统进行合并。我们的方法在保持算法极度紧凑和轻量级的同时,显著提升了以往工作的VPR性能,使其适用于资源受限的平台。