Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch, however is limited by the need for large labelled audiovisual datasets (in each downstream domain of interest). We present AVFormer, a simple method for augmenting audio-only models with visual information, at the same time performing lightweight domain adaptation. We do this by (i) injecting visual embeddings into a frozen ASR model using lightweight trainable adaptors. We show that these can be trained on a small amount of weakly labelled video data with minimum additional training time and parameters. (ii) We also introduce a simple curriculum scheme during training which we show is crucial to enable the model to jointly process audio and visual information effectively; and finally (iii) we show that our model achieves state of the art zero-shot results on three different AV-ASR benchmarks (How2, VisSpeech and Ego4D), while also crucially preserving decent performance on traditional audio-only speech recognition benchmarks (LibriSpeech). Qualitative results show that our model effectively leverages visual information for robust speech recognition.
翻译:摘要:音频-视觉自动语音识别(AV-ASR)旨在通过引入视觉信息提升语音识别系统的鲁棒性。然而,从头开始训练此类全监督多模态模型受限于需要大量带标签的音频-视觉数据集(在每个相关下游领域)。我们提出AVFormer,一种通过视觉信息增强纯音频模型的简单方法,同时实现轻量级领域适配。具体地,我们通过以下方式实现:(i)使用轻量级可训练适配器将视觉嵌入注入冻结的ASR模型。我们证明这种适配器可以在少量弱标签视频数据上训练,且额外训练时间和参数开销极小。(ii)训练过程中引入简单的课程学习策略,实验表明这对使模型能够有效联合处理音频与视觉信息至关重要;(iii)我们的模型在三个不同AV-ASR基准测试(How2、VisSpeech和Ego4D)上取得最先进的零样本结果,同时保持对传统纯音频语音识别基准(LibriSpeech)的优异表现。定性结果表明,该模型能有效利用视觉信息实现鲁棒的语音识别。