Audio-visual person recognition (AVPR) has received extensive attention. However, most datasets used for AVPR research so far are collected in constrained environments, and thus cannot reflect the true performance of AVPR systems in real-world scenarios. To meet the request for research on AVPR in unconstrained conditions, this paper presents a multi-genre AVPR dataset collected `in the wild', named CN-Celeb-AV. This dataset contains more than 420k video segments from 1,136 persons from public media. In particular, we put more emphasis on two real-world complexities: (1) data in multiple genres; (2) segments with partial information. A comprehensive study was conducted to compare CN-Celeb-AV with two popular public AVPR benchmark datasets, and the results demonstrated that CN-Celeb-AV is more in line with real-world scenarios and can be regarded as a new benchmark dataset for AVPR research. The dataset also involves a development set that can be used to boost the performance of AVPR systems in real-life situations. The dataset is free for researchers and can be downloaded from http://cnceleb.org/.
翻译:音视频人物识别(AVPR)近年来受到广泛关注。然而,现有用于AVPR研究的数据集大多在受控环境下采集,无法真实反映AVPR系统在实际场景中的性能。为满足非约束条件下AVPR研究的需求,本文提出一个在自然场景中采集的多体裁AVPR数据集——CN-Celeb-AV。该数据集包含来自1,136位公众人物的超过42万个视频片段。特别地,我们重点关注两种实际场景中的复杂性:(1)多体裁数据;(2)包含不完整信息的片段。通过综合实验将CN-Celeb-AV与两个流行的公开AVPR基准数据集进行比较,结果表明CN-Celeb-AV更贴合真实场景,可作为AVPR研究的基准数据集。该数据集还提供了可用于提升实际场景中AVPR系统性能的开发集。本数据集对研究人员免费开放,可从http://cnceleb.org/下载。