Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have attempted to address this challenge by leveraging separately pre-trained unimodal models to extract visual and textual features. However, these approaches lack the necessary underlying alignment capabilities required to match multimodal data effectively. Besides, these works use prior information to explore explicit part alignments, which may lead to the distortion of intra-modality information. To alleviate these issues, we present IRRA: a cross-modal Implicit Relation Reasoning and Aligning framework that learns relations between local visual-textual tokens and enhances global image-text matching without requiring additional prior supervision. Specifically, we first design an Implicit Relation Reasoning module in a masked language modeling paradigm. This achieves cross-modal interaction by integrating the visual cues into the textual tokens with a cross-modal multimodal interaction encoder. Secondly, to globally align the visual and textual embeddings, Similarity Distribution Matching is proposed to minimize the KL divergence between image-text similarity distributions and the normalized label matching distributions. The proposed method achieves new state-of-the-art results on all three public datasets, with a notable margin of about 3%-9% for Rank-1 accuracy compared to prior methods.
翻译:文本到图像行人检索旨在根据给定的文本描述查询来识别目标行人。其主要挑战在于学习将视觉和文本模态映射到共同的潜在空间中。先前的研究尝试通过利用分别预训练的单模态模型提取视觉和文本特征来解决这一挑战。然而,这些方法缺乏有效匹配多模态数据所需的必要底层对齐能力。此外,这些工作使用先验信息探索显式局部对齐,可能导致模态内信息失真。为解决这些问题,我们提出IRRA:一种跨模态隐式关系推理与对齐框架,该框架学习局部视觉-文本标记间的关系,并增强全局图像-文本匹配,而无需额外先验监督。具体而言,我们首先基于掩码语言建模范式设计了隐式关系推理模块。该模块通过跨模态多模态交互编码器,将视觉线索整合到文本标记中,实现跨模态交互。其次,为全局对齐视觉与文本嵌入,我们提出相似度分布匹配方法,以最小化图像-文本相似度分布与归一化标签匹配分布之间的KL散度。所提方法在三个公开数据集上均取得新的最优结果,与先前方法相比,Rank-1准确率显著提升约3%-9%。