Existing unsupervised person re-identification methods only rely on visual clues to match pedestrians under different cameras. Since visual data is essentially susceptible to occlusion, blur, clothing changes, etc., a promising solution is to introduce heterogeneous data to make up for the defect of visual data. Some works based on full-scene labeling introduce wireless positioning to assist cross-domain person re-identification, but their GPS labeling of entire monitoring scenes is laborious. To this end, we propose to explore unsupervised person re-identification with both visual data and wireless positioning trajectories under weak scene labeling, in which we only need to know the locations of the cameras. Specifically, we propose a novel unsupervised multimodal training framework (UMTF), which models the complementarity of visual data and wireless information. Our UMTF contains a multimodal data association strategy (MMDA) and a multimodal graph neural network (MMGN). MMDA explores potential data associations in unlabeled multimodal data, while MMGN propagates multimodal messages in the video graph based on the adjacency matrix learned from histogram statistics of wireless data. Thanks to the robustness of the wireless data to visual noise and the collaboration of various modules, UMTF is capable of learning a model free of the human label on data. Extensive experimental results conducted on two challenging datasets, i.e., WP-ReID and DukeMTMC-VideoReID demonstrate the effectiveness of the proposed method.
翻译:现有无监督行人再识别方法仅依赖视觉线索在不同摄像头间匹配行人。由于视觉数据本质上易受遮挡、模糊、衣物更换等因素影响,引入异构数据弥补视觉数据缺陷成为可行方案。部分基于全场景标注的工作通过引入无线定位辅助跨域行人再识别,但此类方法需对整个监控场景进行GPS标注,耗费大量人力。为此,我们提出在弱场景标注下同时利用视觉数据与无线定位轨迹进行无监督行人再识别,该方法仅需知晓摄像头位置即可。具体而言,我们设计了一种新型无监督多模态训练框架(UMTF),该框架建模了视觉数据与无线信息的互补性。UMTF包含多模态数据关联策略(MMDA)与多模态图神经网络(MMGN)。MMDA探索未标注多模态数据中的潜在数据关联,而MMGN则基于无线数据直方图统计学习得到的邻接矩阵,在视频图中传播多模态信息。得益于无线数据对视觉噪声的鲁棒性及各模块的协同作用,UMTF能够学习无需人工标注数据的模型。在两个具有挑战性的数据集WP-ReID和DukeMTMC-VideoReID上开展的广泛实验验证了所提方法的有效性。