Discriminative representation is essential to keep a unique identifier for each target in Multiple object tracking (MOT). Some recent MOT methods extract features of the bounding box region or the center point as identity embeddings. However, when targets are occluded, these coarse-grained global representations become unreliable. To this end, we propose exploring diverse fine-grained representation, which describes appearance comprehensively from global and local perspectives. This fine-grained representation requires high feature resolution and precise semantic information. To effectively alleviate the semantic misalignment caused by indiscriminate contextual information aggregation, Flow Alignment FPN (FAFPN) is proposed for multi-scale feature alignment aggregation. It generates semantic flow among feature maps from different resolutions to transform their pixel positions. Furthermore, we present a Multi-head Part Mask Generator (MPMG) to extract fine-grained representation based on the aligned feature maps. Multiple parallel branches of MPMG allow it to focus on different parts of targets to generate local masks without label supervision. The diverse details in target masks facilitate fine-grained representation. Eventually, benefiting from a Shuffle-Group Sampling (SGS) training strategy with positive and negative samples balanced, we achieve state-of-the-art performance on MOT17 and MOT20 test sets. Even on DanceTrack, where the appearance of targets is extremely similar, our method significantly outperforms ByteTrack by 5.0% on HOTA and 5.6% on IDF1. Extensive experiments have proved that diverse fine-grained representation makes Re-ID great again in MOT.
翻译:判别性表征对于在多目标跟踪中保持每个目标的唯一标识符至关重要。近期的一些多目标跟踪方法提取边界框区域或中心点的特征作为身份嵌入。然而,当目标被遮挡时,这些粗粒度的全局表征变得不可靠。为此,我们提出探索多样细粒度表征,从全局和局部角度全面描述外观。这种细粒度表征需要高特征分辨率和精确的语义信息。为有效缓解由无差别上下文信息聚合导致的语义错位问题,我们提出流对齐特征金字塔网络用于多尺度特征对齐聚合,通过在不同分辨率的特征图间生成语义流来变换像素位置。此外,我们提出多头部分掩码生成器,基于对齐的特征图提取细粒度表征。MPMG的多个并行分支使其能够聚焦于目标的不同部分,在无需标签监督的情况下生成局部掩码。目标掩码中的多样细节有助于细粒度表征。最终,通过正负样本平衡的洗牌分组采样训练策略,我们在MOT17和MOT20测试集上达到了先进性能。即使在目标外观高度相似的DanceTrack数据集上,我们的方法在HOTA和IDF1指标上分别以5.0%和5.6%的显著优势超越ByteTrack。大量实验证明,多样细粒度表征使多目标跟踪中的重识别再次焕发活力。