Audio-visual automatic speech recognition (AV-ASR) models are very effective at reducing word error rates on noisy speech, but require large amounts of transcribed AV training data. Recently, audio-visual self-supervised learning (SSL) approaches have been developed to reduce this dependence on transcribed AV data, but these methods are quite complex and computationally expensive. In this work, we propose replacing these expensive AV-SSL methods with a simple and fast \textit{audio-only} SSL method, and then performing AV supervised fine-tuning. We show that this approach is competitive with state-of-the-art (SOTA) AV-SSL methods on the LRS3-TED benchmark task (within 0.5% absolute WER), while being dramatically simpler and more efficient (12-30x faster to pre-train). Furthermore, we show we can extend this approach to convert a SOTA audio-only ASR model into an AV model. By doing so, we match SOTA AV-SSL results, even though no AV data was used during pre-training.
翻译:音频-视觉自动语音识别(AV-ASR)模型在降低含噪语音的词错误率方面非常有效,但需要大量带标注的音频-视觉训练数据。近年来,为减少对标注音频-视觉数据的依赖,研究人员开发了音频-视觉自监督学习(SSL)方法,但这类方法复杂度高且计算成本昂贵。本研究提出用简单快速的纯音频自监督学习方法替代成本高昂的音频-视觉自监督学习方法,随后进行音频-视觉监督微调。我们证明,该方法在LRS3-TED基准任务上能与当前最优音频-视觉自监督学习方法竞争(绝对词错误率差异在0.5%以内),且显著更简单高效(预训练速度快12-30倍)。此外,我们验证了该方法可将当前最优的纯音频自动语音识别模型扩展为音频-视觉模型。通过此方式,即使预训练阶段未使用任何音频-视觉数据,我们仍能达到与当前最优音频-视觉自监督学习方法相匹配的结果。