Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID, TCD-TIMIT, and Lip2Wav, can have data asynchrony issues. Training lip-to-speech with such datasets may further cause the model asynchrony issue -- that is, the generated speech and the input video are out of sync. To address these asynchrony issues, we propose a synchronized lip-to-speech (SLTS) model with an automatic synchronization mechanism (ASM) to correct data asynchrony and penalize model asynchrony. We further demonstrate the limitation of the commonly adopted evaluation metrics for LTS with asynchronous test data and introduce an audio alignment frontend before the metrics sensitive to time alignment for better evaluation. We compare our method with state-of-the-art approaches on conventional and time-aligned metrics to show the benefits of synchronization training.
翻译:大多数唇语到语音(LTS)合成模型均基于数据集中的音视频对完美同步这一假设进行训练与评估。本研究表明,常用的音视频数据集(如GRID、TCD-TIMIT和Lip2Wav)可能存在数据异步问题。使用此类数据集训练唇语到语音模型会进一步引发模型异步问题——即生成的语音与输入视频不同步。为解决这些异步问题,我们提出一种带有自动同步机制(ASM)的同步唇语到语音(SLTS)模型,用于修正数据异步并惩罚模型异步。我们进一步论证了针对异步测试数据的LTS常用评估指标存在的局限性,并提出在时序对齐敏感的评估指标前引入音频对齐前端以优化评估效果。通过与传统方法及时间对齐指标的对比实验,我们展示了同步训练方法的优势。