Recent researches on unsupervised person re-identification~(reID) have demonstrated that pre-training on unlabeled person images achieves superior performance on downstream reID tasks than pre-training on ImageNet. However, those pre-trained methods are specifically designed for reID and suffer flexible adaption to other pedestrian analysis tasks. In this paper, we propose VAL-PAT, a novel framework that learns transferable representations to enhance various pedestrian analysis tasks with multimodal information. To train our framework, we introduce three learning objectives, \emph{i.e.,} self-supervised contrastive learning, image-text contrastive learning and multi-attribute classification. The self-supervised contrastive learning facilitates the learning of the intrinsic pedestrian properties, while the image-text contrastive learning guides the model to focus on the appearance information of pedestrians.Meanwhile, multi-attribute classification encourages the model to recognize attributes to excavate fine-grained pedestrian information. We first perform pre-training on LUPerson-TA dataset, where each image contains text and attribute annotations, and then transfer the learned representations to various downstream tasks, including person reID, person attribute recognition and text-based person search. Extensive experiments demonstrate that our framework facilitates the learning of general pedestrian representations and thus leads to promising results on various pedestrian analysis tasks.
翻译:近年来,无监督行人重识别(reID)研究表明,在无标签行人图像上预训练相比在ImageNet上预训练,能在下游reID任务中取得更优性能。然而,这些预训练方法是专为reID设计的,难以灵活适配其他行人分析任务。本文提出VAL-PAT框架,这是一种基于多模态信息学习可迁移表征以增强多种行人分析任务的新型架构。为训练该框架,我们引入三个学习目标:自监督对比学习、图像-文本对比学习和多属性分类。自监督对比学习促进行人内在属性的学习,而图像-文本对比学习引导模型聚焦行人的外观信息。同时,多属性分类鼓励模型识别属性以挖掘细粒度行人信息。我们首先在LUPerson-TA数据集上进行预训练(该数据集中每张图像均包含文本和属性标注),然后将学习到的表征迁移至行人重识别、行人属性识别和基于文本的行人搜索等多种下游任务。大量实验表明,该框架有助于学习通用的行人表征,从而在各类行人分析任务中取得优异结果。