Text-based Person Retrieval aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited large-scale datasets. In this paper, we introduce a CLIP-based Synergistic Knowledge Transfer(CSKT) approach for TBPR. Specifically, to explore the CLIP's knowledge on input side, we first propose a Bidirectional Prompts Transferring (BPT) module constructed by text-to-image and image-to-text bidirectional prompts and coupling projections. Secondly, Dual Adapters Transferring (DAT) is designed to transfer knowledge on output side of Multi-Head Self-Attention (MHSA) in vision and language. This synergistic two-way collaborative mechanism promotes the early-stage feature fusion and efficiently exploits the existing knowledge of CLIP. CSKT outperforms the state-of-the-art approaches across three benchmark datasets when the training parameters merely account for 7.4% of the entire model, demonstrating its remarkable efficiency, effectiveness and generalization.
翻译:文本行人检索旨在根据文本查询检索目标行人图像。其主要挑战在于弥合视觉与语言模态之间的巨大鸿沟,尤其是在大规模数据集有限的情况下。本文提出一种基于CLIP的协同知识迁移方法用于文本行人检索。具体而言,为探索CLIP在输入侧的知识,我们首先设计了一种双向提示迁移模块,该模块由文本到图像和图像到文本的双向提示以及耦合投影构成。其次,我们设计了双重适配器迁移模块,用于迁移视觉和语言中多头自注意力输出侧的知识。这种协同双向协作机制促进了早期特征融合,有效利用了CLIP的既有知识。当训练参数仅占整个模型的7.4%时,CSKT在三个基准数据集上均超越了现有最优方法,展现出卓越的效率、有效性和泛化能力。