Generating high-quality and person-generic visual dubbing remains a challenge. Recent innovation has seen the advent of a two-stage paradigm, decoupling the rendering and lip synchronization process facilitated by intermediate representation as a conduit. Still, previous methodologies rely on rough landmarks or are confined to a single speaker, thus limiting their performance. In this paper, we propose DiffDub: Diffusion-based dubbing. We first craft the Diffusion auto-encoder by an inpainting renderer incorporating a mask to delineate editable zones and unaltered regions. This allows for seamless filling of the lower-face region while preserving the remaining parts. Throughout our experiments, we encountered several challenges. Primarily, the semantic encoder lacks robustness, constricting its ability to capture high-level features. Besides, the modeling ignored facial positioning, causing mouth or nose jitters across frames. To tackle these issues, we employ versatile strategies, including data augmentation and supplementary eye guidance. Moreover, we encapsulated a conformer-based reference encoder and motion generator fortified by a cross-attention mechanism. This enables our model to learn person-specific textures with varying references and reduces reliance on paired audio-visual data. Our rigorous experiments comprehensively highlight that our ground-breaking approach outpaces existing methods with considerable margins and delivers seamless, intelligible videos in person-generic and multilingual scenarios.
翻译:摘要:生成高质量且通用的人物视觉配音仍是一项挑战。近期创新催生了一种两阶段范式,通过中间表示作为桥梁,将渲染与唇同步过程解耦。然而,先前方法依赖粗略的面部特征点,或局限于单一说话者,从而限制了其性能。本文提出DiffDub:基于扩散的配音。我们首先通过修补渲染器构建扩散自编码器,该渲染器引入掩码以区分可编辑区域与未修改区域,从而在保留其余部分的同时无缝填充下半脸区域。在实验中,我们遇到了若干挑战。首先,语义编码器缺乏鲁棒性,限制了其捕捉高层特征的能力。此外,建模忽略了面部定位,导致帧间嘴巴或鼻子抖动。为解决这些问题,我们采用了多种策略,包括数据增强与辅助眼睛引导。同时,我们封装了基于Conformer的参考编码器与由交叉注意力机制强化的运动生成器。这使得模型能够学习具有不同参考图像的人物特定纹理,并减少对配对音视频数据的依赖。通过严格实验,我们全面展示了这一突破性方法以显著优势超越现有方法,并在通用人物与多语言场景下生成流畅、清晰的视频。