Modern generators render talking-head videos with impressive photorealism, ushering in new user experiences such as videoconferencing under constrained bandwidth budgets. Their safe adoption, however, requires a mechanism to verify if the rendered video is trustworthy. For instance, for videoconferencing we must identify cases in which a synthetic video portrait uses the appearance of an individual without their consent. We term this task avatar fingerprinting. Specifically, we learn an embedding in which the motion signatures of one identity are grouped together, and pushed away from those of the other identities. This allows us to link the synthetic video to the identity driving the expressions in the video, regardless of the facial appearance shown. Avatar fingerprinting algorithms will be critical as talking head generators become more ubiquitous, and yet no large scale datasets exist for this new task. Therefore, we contribute a large dataset of people delivering scripted and improvised short monologues, accompanied by synthetic videos in which we render videos of one person using the facial appearance of another. Project page: https://research.nvidia.com/labs/nxp/avatar-fingerprinting/.
翻译:现代生成器能够以惊人的照片真实感渲染说话人头视频,从而在带宽受限的情况下开启视频会议等全新用户体验。然而,这类技术的安全应用需要建立机制验证生成视频的可信度。例如,在视频会议场景中,我们必须识别出合成视频肖像是否存在未经同意使用他人外貌特征的情况。我们将此任务命名为"身份指纹识别"。具体而言,我们学习一个嵌入空间,其中同一身份的运动特征被聚类在一起,同时与其他身份的运动特征相分离。这使得我们能够将合成视频与驱动视频表情的身份关联起来,而无需依赖视频中展示的面部外貌。随着说话人头生成技术的日益普及,身份指纹识别算法将变得至关重要,但目前针对这一新任务尚缺乏大规模数据集。为此,我们贡献了一个大型数据集,包含人们按脚本和即兴表演的短篇独白,以及配套的合成视频——这些视频使用一人外貌特征渲染另一人的面部。项目页面:https://research.nvidia.com/labs/nxp/avatar-fingerprinting/。