In this paper, we propose a novel network framework for indoor 3D object detection to handle variable input frame numbers in practical scenarios. Existing methods only consider fixed frames of input data for a single detector, such as monocular RGB-D images or point clouds reconstructed from dense multi-view RGB-D images. While in practical application scenes such as robot navigation and manipulation, the raw input to the 3D detectors is the RGB-D images with variable frame numbers instead of the reconstructed scene point cloud. However, the previous approaches can only handle fixed frame input data and have poor performance with variable frame input. In order to facilitate 3D object detection methods suitable for practical tasks, we present a novel 3D detection framework named AnyView for our practical applications, which generalizes well across different numbers of input frames with a single model. To be specific, we propose a geometric learner to mine the local geometric features of each input RGB-D image frame and implement local-global feature interaction through a designed spatial mixture module. Meanwhile, we further utilize a dynamic token strategy to adaptively adjust the number of extracted features for each frame, which ensures consistent global feature density and further enhances the generalization after fusion. Extensive experiments on the ScanNet dataset show our method achieves both great generalizability and high detection accuracy with a simple and clean architecture containing a similar amount of parameters with the baselines.
翻译:本文提出一种面向室内三维目标检测的新型网络框架,旨在处理实际场景中可变输入帧数的问题。现有方法仅考虑单一检测器对固定帧数输入数据的处理能力,例如单目RGB-D图像或由密集多视角RGB-D图像重建的点云。但在机器人导航与操作等实际应用场景中,三维检测器的原始输入是帧数可变的RGB-D图像,而非重建后的场景点云。然而,现有方法仅能处理固定帧输入数据,在可变帧输入场景中性能表现不佳。为促进适用于实际任务的三维目标检测方法发展,我们提出一种名为AnyView的新型三维检测框架,该框架通过单一模型即可在不同输入帧数下实现良好泛化。具体而言,我们设计几何学习器挖掘每个输入RGB-D图像帧的局部几何特征,并通过所设计的空间混合模块实现局部-全局特征交互。同时,进一步采用动态标记策略自适应调整每帧提取的特征数量,确保融合后全局特征密度的一致性,从而增强泛化能力。在ScanNet数据集上的大量实验表明,本方法在保持与基线模型相似参数量的简洁架构下,兼具优异的泛化性能与高检测精度。