In this paper, we propose a novel Lip-to-Speech synthesis (L2S) framework, for synthesizing intelligible speech from a silent lip movement video. Specifically, to complement the insufficient supervisory signal of the previous L2S model, we propose to use quantized self-supervised speech representations, named speech units, as an additional prediction target for the L2S model. Therefore, the proposed L2S model is trained to generate multiple targets, mel-spectrogram and speech units. As the speech units are discrete while mel-spectrogram is continuous, the proposed multi-target L2S model can be trained with strong content supervision, without using text-labeled data. Moreover, to accurately convert the synthesized mel-spectrogram into a waveform, we introduce a multi-input vocoder that can generate a clear waveform even from blurry and noisy mel-spectrogram by referring to the speech units. Extensive experimental results confirm the effectiveness of the proposed method in L2S.
翻译:本文提出了一种新颖的唇语到语音合成(L2S)框架,旨在从无声嘴唇运动视频中合成可理解的语音。具体而言,为弥补先前L2S模型监督信号不足的问题,我们提出使用量化的自监督语音表征(称为语音单元)作为L2S模型的额外预测目标。因此,所提出的L2S模型被训练用于生成多个目标——梅尔频谱图和语音单元。由于语音单元是离散的而梅尔频谱图是连续的,所提出的多目标L2S模型可在无需文本标注数据的情况下,通过强内容监督进行训练。此外,为将合成的梅尔频谱图准确转换为波形,我们引入了一种多输入声码器,即使面对模糊且含噪的梅尔频谱图,也能通过参考语音单元生成清晰波形。大量实验结果验证了所提方法在L2S中的有效性。