Traditional text-based person re-identification (ReID) techniques heavily rely on fully matched multi-modal data, which is an ideal scenario. However, due to inevitable data missing and corruption during the collection and processing of cross-modal data, the incomplete data issue is usually met in real-world applications. Therefore, we consider a more practical task termed the incomplete text-based ReID task, where person images and text descriptions are not completely matched and contain partially missing modality data. To this end, we propose a novel Prototype-guided Cross-modal Completion and Alignment (PCCA) framework to handle the aforementioned issues for incomplete text-based ReID. Specifically, we cannot directly retrieve person images based on a text query on missing modality data. Therefore, we propose the cross-modal nearest neighbor construction strategy for missing data by computing the cross-modal similarity between existing images and texts, which provides key guidance for the completion of missing modal features. Furthermore, to efficiently complete the missing modal features, we construct the relation graphs with the aforementioned cross-modal nearest neighbor sets of missing modal data and the corresponding prototypes, which can further enhance the generated missing modal features. Additionally, for tighter fine-grained alignment between images and texts, we raise a prototype-aware cross-modal alignment loss that can effectively reduce the modality heterogeneity gap for better fine-grained alignment in common space. Extensive experimental results on several benchmarks with different missing ratios amply demonstrate that our method can consistently outperform state-of-the-art text-image ReID approaches.
翻译:传统基于文本的行人重识别(ReID)技术严重依赖完全匹配的多模态数据,这仅是一种理想场景。然而,由于跨模态数据采集与处理过程中不可避免的数据缺失与损坏,实际应用中常面临不完整数据问题。因此,我们考虑一个更具实践性的任务——不完整文本行人重识别,该任务中行人图像与文本描述并非完全匹配,且存在部分模态数据缺失。为此,我们提出一种新型原型引导的跨模态补全与对齐(PCCA)框架,以解决上述不完整文本行人重识别问题。具体而言,在缺失模态数据情况下,我们无法直接基于文本查询检索行人图像。因此,我们提出一种面向缺失数据的跨模态最近邻构建策略,通过计算现有图像与文本间的跨模态相似度,为缺失模态特征的补全提供关键引导。进一步地,为高效补全缺失模态特征,我们利用上述缺失模态数据的跨模态最近邻集合及对应原型构建关系图,从而增强生成的缺失模态特征。此外,为实现图像与文本间更紧密的细粒度对齐,我们提出一种原型感知的跨模态对齐损失函数,该函数可有效降低模态异质性差距,从而在公共空间中实现更优的细粒度对齐。在多个不同缺失比例基准数据集上的大量实验结果表明,我们的方法能够持续超越当前最先进的文本-图像行人重识别方法。