Person clustering with multi-modal clues, including faces, bodies, and voices, is critical for various tasks, such as movie parsing and identity-based movie editing. Related methods such as multi-view clustering mainly project multi-modal features into a joint feature space. However, multi-modal clue features are usually rather weakly correlated due to the semantic gap from the modality-specific uniqueness. As a result, these methods are not suitable for person clustering. In this paper, we propose a Relation-Aware Distribution representation Network (RAD-Net) to generate a distribution representation for multi-modal clues. The distribution representation of a clue is a vector consisting of the relation between this clue and all other clues from all modalities, thus being modality agnostic and good for person clustering. Accordingly, we introduce a graph-based method to construct distribution representation and employ a cyclic update policy to refine distribution representation progressively. Our method achieves substantial improvements of +6% and +8.2% in F-score on the Video Person-Clustering Dataset (VPCD) and VoxCeleb2 multi-view clustering dataset, respectively. Codes will be released publicly upon acceptance.
翻译:多模态线索的人物聚类(包括人脸、身体和声音)对电影解析、基于身份的电影剪辑等任务至关重要。多视图聚类等相关方法主要将多模态特征投影到联合特征空间中。然而,由于模态特有属性造成的语义鸿沟,多模态线索特征之间通常相关性较弱,导致此类方法不适用于人物聚类。本文提出关系感知式分布表示网络(RAD-Net),为多模态线索生成分布表示。每条线索的分布表示由该线索与所有其他模态线索的关系向量构成,因此具有模态无关性,适用于人物聚类。为此,我们引入基于图的方法构建分布表示,并采用循环更新策略逐步优化分布表示。在视频人物聚类数据集(VPCD)和VoxCeleb2多视图聚类数据集上,本方法的F值分别提升6%和8.2%。代码将在论文接收后公开。