Recent deep learning-based multi-view people detection (MVD) methods have shown promising results on existing datasets. However, current methods are mainly trained and evaluated on small, single scenes with a limited number of multi-view frames and fixed camera views. As a result, these methods may not be practical for detecting people in larger, more complex scenes with severe occlusions and camera calibration errors. This paper focuses on improving multi-view people detection by developing a supervised view-wise contribution weighting approach that better fuses multi-camera information under large scenes. Besides, a large synthetic dataset is adopted to enhance the model's generalization ability and enable more practical evaluation and comparison. The model's performance on new testing scenes is further improved with a simple domain adaptation technique. Experimental results demonstrate the effectiveness of our approach in achieving promising cross-scene multi-view people detection performance. See code here: https://vcc.tech/research/2024/MVD.
翻译:近年来,基于深度学习的多视角行人检测方法在现有数据集上取得了良好的效果。然而,当前方法主要在规模较小、视角单一的场景中进行训练和评估,这些场景的多视角帧数量有限且相机视角固定。因此,这些方法可能不适用于在遮挡严重、存在相机标定误差的更大、更复杂的场景中检测行人。本文通过开发一种监督视角贡献加权方法,以改善多视角行人检测,该方法能在大场景下更好地融合多相机信息。此外,采用大规模合成数据集以增强模型的泛化能力,并实现更具实践意义的评估与比较。通过简单的域适应技术,模型在新测试场景上的性能得到进一步提升。实验结果证明了我们的方法在实现良好跨场景多视角行人检测性能方面的有效性。代码参见:https://vcc.tech/research/2024/MVD。