In this paper, we study the problem of semi-supervised 3D object detection, which is of great importance considering the high annotation cost for cluttered 3D indoor scenes. We resort to the robust and principled framework of selfteaching, which has triggered notable progress for semisupervised learning recently. While this paradigm is natural for image-level or pixel-level prediction, adapting it to the detection problem is challenged by the issue of proposal matching. Prior methods are based upon two-stage pipelines, matching heuristically selected proposals generated in the first stage and resulting in spatially sparse training signals. In contrast, we propose the first semisupervised 3D detection algorithm that works in the singlestage manner and allows spatially dense training signals. A fundamental issue of this new design is the quantization error caused by point-to-voxel discretization, which inevitably leads to misalignment between two transformed views in the voxel domain. To this end, we derive and implement closed-form rules that compensate this misalignment onthe-fly. Our results are significant, e.g., promoting ScanNet [email protected] from 35.2% to 48.5% using 20% annotation. Codes and data will be publicly available.
翻译:本文研究了半监督3D目标检测问题,鉴于杂乱3D室内场景的高标注成本,该问题具有重要价值。我们采用稳健且原理化的自教学框架,该方法近期在半监督学习中取得了显著进展。尽管该范式自然适用于图像级或像素级预测,但将其适配至检测问题面临候选框匹配的挑战。现有方法基于两阶段流程,匹配第一阶段启发式选择的候选框,导致空间稀疏的训练信号。相比之下,我们提出首个单阶段半监督3D检测算法,能够实现空间密集的训练信号。该新设计的一个根本问题是点体素离散化导致的量化误差,该误差不可避免地在体素域中造成两种变换视图之间的错位。为此,我们推导并实现了能够在线补偿这种错位的闭式规则。我们的成果显著,例如在仅使用20%标注数据的情况下,将ScanNet的[email protected]从35.2%提升至48.5%。代码与数据将公开提供。