Detection-based methods have been viewed unfavorably in crowd analysis due to their poor performance in dense crowds. However, we argue that the potential of these methods has been underestimated, as they offer crucial information for crowd analysis that is often ignored. Specifically, the area size and confidence score of output proposals and bounding boxes provide insight into the scale and density of the crowd. To leverage these underutilized features, we propose Crowd Hat, a plug-and-play module that can be easily integrated with existing detection models. This module uses a mixed 2D-1D compression technique to refine the output features and obtain the spatial and numerical distribution of crowd-specific information. Based on these features, we further propose region-adaptive NMS thresholds and a decouple-then-align paradigm that address the major limitations of detection-based methods. Our extensive evaluations on various crowd analysis tasks, including crowd counting, localization, and detection, demonstrate the effectiveness of utilizing output features and the potential of detection-based methods in crowd analysis.
翻译:基于检测的方法因在密集人群场景中表现不佳而在人群分析中不受青睐。然而,我们认为这些方法的潜力被低估了,因为它们提供了常被忽视的人群分析关键信息。具体而言,输出候选框和边界框的面积大小与置信度分数可揭示人群的尺度与密度分布。为充分利用这些未开发的特性,我们提出人群帽子模块——一种即插即用模块,可轻松集成到现有检测模型中。该模块采用混合二维-一维压缩技术对输出特征进行精炼,获取人群特定信息的空间与数值分布。基于这些特征,我们进一步提出区域自适应非极大值抑制阈值和解耦-对齐范式,有效克服了基于检测方法的主要局限性。我们在包括人群计数、定位和检测等多种人群分析任务上的广泛评估表明,利用输出特征具有显著效果,并且基于检测的方法在人群分析中具有巨大潜力。