Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head video clips which simultaneously have accurate lip synchronization and motion smoothness. Previous approaches, including 3DMM-based (3D Morphable Model) methods and NeRF-based (Neural Radiance Field) methods, are sub-optimal in that they either require minutes of source videos and days of training time or lack the disentangled control of verbal (e.g., lip motion) and non-verbal (e.g., head pose and expression) representations for video clip insertion. In this work, we fully utilize the video context to design a novel framework for talking-head video editing, which achieves efficiency, disentangled motion control, and sequential smoothness. Specifically, we decompose this framework to motion prediction and motion-conditioned rendering: (1) We first design an animation prediction module that efficiently obtains smooth and lip-sync motion sequences conditioned on the driven speech. This module adopts a non-autoregressive network to obtain context prior and improve the prediction efficiency, and it learns a speech-animation mapping prior with better generalization to novel speech from a multi-identity video dataset. (2) We then introduce a neural rendering module to synthesize the photo-realistic and full-head video frames given the predicted motion sequence. This module adopts a pre-trained head topology and uses only few frames for efficient fine-tuning to obtain a person-specific rendering model. Extensive experiments demonstrate that our method efficiently achieves smoother editing results with higher image quality and lip accuracy using less data than previous methods.
翻译:说话人头部视频编辑旨在通过文本转录编辑器高效地插入、删除和替换预先录制视频中的词语。该任务的关键挑战是获得一个编辑模型,能够生成同时具备精确唇同步和运动平滑性的新说话人头部视频片段。先前的方法,包括基于3D可变形模型(3DMM)的方法和基于神经辐射场(NeRF)的方法,都非最优:它们要么需要数分钟的源视频和数天的训练时间,要么缺乏对言语(如唇部运动)和非言语(如头部姿态和表情)表征的解耦控制,无法用于视频片段插入。在本工作中,我们充分利用视频上下文设计了一个用于说话人头部视频编辑的新型框架,实现了高效性、解耦运动控制以及序列平滑性。具体而言,我们将该框架分解为运动预测和运动条件渲染两个部分:(1)首先,我们设计了一个动画预测模块,该模块在驱动语音条件下高效地获得平滑且唇同步的运动序列。该模块采用非自回归网络来获取上下文先验并提高预测效率,并且从多身份视频数据集中学习一个具有更好泛化能力的语音-动画映射先验,以适应新颖语音。(2)然后,我们引入一个神经渲染模块,在给定预测运动序列的情况下合成逼真的全头部视频帧。该模块采用预训练的头部拓扑结构,仅利用少量帧进行高效微调,以获得特定人物的渲染模型。大量实验表明,与先前方法相比,我们的方法在利用更少数据的情况下,能够高效地实现更平滑的编辑结果,并具有更高的图像质量和唇部准确性。