Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a high-quality large-scale singing head dataset, SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we argue that 3D and 2D facial animation tasks can be solved together, and propose a unified singing facial animation framework named UniSinger to achieve both singing audio-driven 3D singing head animation and 2D singing portrait video synthesis. Extensive comparative experiments with both SOTA 3D facial animation and 2D portrait animation methods demonstrate the necessity of singing-specific datasets in singing head animation tasks and the promising performance of our unified facial animation framework.
翻译:歌唱作为仅次于说话常见面部动作,可视为跨越种族与文化的通用语言,在情感交流、艺术和娱乐中具有重要作用。然而,由于缺乏歌唱头部数据集,且歌唱与说话在节奏和幅度上存在领域差异,该领域在音频驱动面部动画研究中常被忽视。为此,我们收集了高质量大规模歌唱头部数据集SingingHead,包含来自76名个体和8种音乐类型的超过27小时的同步歌唱视频、3D面部运动、歌唱音频及背景音乐。基于SingingHead数据集,我们提出3D与2D面部动画任务可协同解决的观点,并设计了统一歌唱面部动画框架UniSinger,可同时实现歌唱音频驱动的3D歌唱头部动画与2D歌唱肖像视频合成。与现有最先进的3D面部动画和2D肖像动画方法的广泛对比实验表明,歌唱专用数据集在歌唱头部动画任务中的必要性,以及我们统一面部动画框架的优越性能。