Striking a balance between precision and efficiency presents a prominent challenge in the bird's-eye-view (BEV) 3D object detection. Although previous camera-based BEV methods achieved remarkable performance by incorporating long-term temporal information, most of them still face the problem of low efficiency. One potential solution is knowledge distillation. Existing distillation methods only focus on reconstructing spatial features, while overlooking temporal knowledge. To this end, we propose TempDistiller, a Temporal knowledge Distiller, to acquire long-term memory from a teacher detector when provided with a limited number of frames. Specifically, a reconstruction target is formulated by integrating long-term temporal knowledge through self-attention operation applied to feature teachers. Subsequently, novel features are generated for masked student features via a generator. Ultimately, we utilize this reconstruction target to reconstruct the student features. In addition, we also explore temporal relational knowledge when inputting full frames for the student model. We verify the effectiveness of the proposed method on the nuScenes benchmark. The experimental results show our method obtain an enhancement of +1.6 mAP and +1.1 NDS compared to the baseline, a speed improvement of approximately 6 FPS after compressing temporal knowledge, and the most accurate velocity estimation.
翻译:在鸟瞰视角下的三维物体检测中,平衡精度与效率是一个突出挑战。尽管以往的基于摄像头的鸟瞰图方法通过整合长期时间信息取得了显著性能,但大多数方法仍面临效率低下的问题。知识蒸馏是一种潜在解决方案。现有蒸馏方法仅关注重建空间特征,而忽略了时间知识。为此,我们提出TempDistiller——一种时间知识蒸馏器,能在有限帧输入下从教师检测器获取长期记忆。具体地,通过对特征教师应用自注意力操作整合长期时间知识,构建重建目标。随后,通过生成器为掩码学生特征生成新特征。最终,利用该重建目标对学生特征进行重建。此外,我们还探索了输入完整帧时学生模型的时间关系知识。我们在nuScenes基准上验证了所提方法的有效性。实验结果表明,与基线相比,我们的方法获得了+1.6 mAP和+1.1 NDS的提升,压缩时间知识后速度提升约6 FPS,并实现了最准确的速度估计。