In this paper, we study the problem of semi-supervised 3D object detection, which is of great importance considering the high annotation cost for cluttered 3D indoor scenes. We resort to the robust and principled framework of selfteaching, which has triggered notable progress for semisupervised learning recently. While this paradigm is natural for image-level or pixel-level prediction, adapting it to the detection problem is challenged by the issue of proposal matching. Prior methods are based upon two-stage pipelines, matching heuristically selected proposals generated in the first stage and resulting in spatially sparse training signals. In contrast, we propose the first semisupervised 3D detection algorithm that works in the singlestage manner and allows spatially dense training signals. A fundamental issue of this new design is the quantization error caused by point-to-voxel discretization, which inevitably leads to misalignment between two transformed views in the voxel domain. To this end, we derive and implement closed-form rules that compensate this misalignment onthe-fly. Our results are significant, e.g., promoting ScanNet [email protected] from 35.2% to 48.5% using 20% annotation. Codes and data will be publicly available.
翻译:本文研究了半监督3D物体检测问题,考虑到杂乱的3D室内场景标注成本高昂,该问题具有重要意义。我们采用了鲁棒且有原则的自教学框架,该框架最近在推动半监督学习方面取得了显著进展。尽管这一范式自然适用于图像级或像素级预测,但将其适配到检测问题受到提案匹配挑战的制约。现有方法基于两阶段流程,匹配第一阶段启发式选择的提案,导致空间稀疏的训练信号。相比之下,我们提出了首个以单阶段方式运行并允许空间密集训练信号的半监督3D检测算法。这一新设计的基本问题在于点-体素离散化引起的量化误差,这不可避免地导致体素域中两个变换视角之间的错位。为此,我们推导并实现了在线补偿这种错位的闭式规则。我们的结果显著,例如使用20%标注将ScanNet的[email protected]从35.2%提升至48.5%。代码和数据将公开提供。