Visible-infrared person re-identification (VI-ReID) aims to match specific pedestrian images from different modalities. Although suffering an extra modality discrepancy, existing methods still follow the softmax loss training paradigm, which is widely used in single-modality classification tasks. The softmax loss lacks an explicit penalty for the apparent modality gap, which adversely limits the performance upper bound of the VI-ReID task. In this paper, we propose the spectral-aware softmax (SA-Softmax) loss, which can fully explore the embedding space with the modality information and has clear interpretability. Specifically, SA-Softmax loss utilizes an asynchronous optimization strategy based on the modality prototype instead of the synchronous optimization based on the identity prototype in the original softmax loss. To encourage a high overlapping between two modalities, SA-Softmax optimizes each sample by the prototype from another spectrum. Based on the observation and analysis of SA-Softmax, we modify the SA-Softmax with the Feature Mask and Absolute-Similarity Term to alleviate the ambiguous optimization during model training. Extensive experimental evaluations conducted on RegDB and SYSU-MM01 demonstrate the superior performance of the SA-Softmax over the state-of-the-art methods in such a cross-modality condition.
翻译:可见光-红外行人重识别(VI-ReID)旨在匹配不同模态下的特定行人图像。尽管面临额外的模态差异问题,现有方法仍沿用广泛用于单模态分类任务的Softmax损失函数训练范式。由于缺乏对显式模态差异的惩罚机制,该损失函数从根本上限制了VI-ReID任务的性能上限。本文提出频谱感知Softmax损失函数(SA-Softmax),该函数能充分利用模态信息探索特征嵌入空间,并具有清晰的可解释性。具体而言,SA-Softmax采用基于模态原型的异步优化策略,替代原始Softmax损失中基于身份原型的同步优化。为促进两种模态间的高重叠性,该损失函数通过另一频谱的原型对各样本进行优化。基于对SA-Softmax的观察与分析,我们引入特征掩膜与绝对相似度项对其进行改进,以缓解模型训练中的模糊优化问题。在RegDB和SYSU-MM01数据集上的大量实验评估表明,SA-Softmax在跨模态场景下展现了优于现有最先进方法的性能。