Visible-Infrared person Re-IDentification (VI-ReID) is a challenging cross-modality image retrieval task that aims to match pedestrians' images across visible and infrared cameras. To solve the modality gap, existing mainstream methods adopt a learning paradigm converting the image retrieval task into an image classification task with cross-entropy loss and auxiliary metric learning losses. These losses follow the strategy of adjusting the distribution of extracted embeddings to reduce the intra-class distance and increase the inter-class distance. However, such objectives do not precisely correspond to the final test setting of the retrieval task, resulting in a new gap at the optimization level. By rethinking these keys of VI-ReID, we propose a simple and effective method, the Multi-level Cross-modality Joint Alignment (MCJA), bridging both modality and objective-level gap. For the former, we design the Modality Alignment Augmentation, which consists of three novel strategies, the weighted grayscale, cross-channel cutmix, and spectrum jitter augmentation, effectively reducing modality discrepancy in the image space. For the latter, we introduce a new Cross-Modality Retrieval loss. It is the first work to constrain from the perspective of the ranking list, aligning with the goal of the testing stage. Moreover, based on the global feature only, our method exhibits good performance and can serve as a strong baseline method for the VI-ReID community.
翻译:可见光-红外行人重识别(VI-ReID)是一项具有挑战性的跨模态图像检索任务,旨在匹配可见光与红外相机下的行人图像。为消除模态差异,现有主流方法采用将图像检索任务转化为图像分类任务的学习范式,结合交叉熵损失与辅助度量学习损失。这些损失函数遵循调整嵌入特征分布的策略,以减小类内距离、增大类间距离。然而,此类优化目标与检索任务的最终测试设定并不完全一致,由此在优化层面产生了新的鸿沟。通过重新审视VI-ReID的关键问题,我们提出了一种简洁高效的方法——多层级跨模态联合对齐(MCJA),同时弥合了模态层级与目标层级的鸿沟。针对前者,我们设计了模态对齐增强模块,包含三种创新策略:权重灰度化、跨通道CutMix和频谱抖动增强,有效降低了图像空间的模态差异。针对后者,我们引入了新的跨模态检索损失函数。这是首个从排序列表视角进行约束的工作,与测试阶段目标高度一致。此外,仅基于全局特征,我们的方法便展现出优异性能,可成为VI-ReID社区的强基准方法。