Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations. However, such representations are prone to over-reliance on audio modality as it is much easier to recognize than video modality in clean conditions. As a result, the AVSR model underestimates the importance of visual stream in face of noise corruption. To this end, we leverage visual modality-specific representations to provide stable complementary information for the AVSR task. Specifically, we propose a reinforcement learning (RL) based framework called MSRL, where the agent dynamically harmonizes modality-invariant and modality-specific representations in the auto-regressive decoding process. We customize a reward function directly related to task-specific metrics (i.e., word error rate), which encourages the MSRL to effectively explore the optimal integration strategy. Experimental results on the LRS3 dataset show that the proposed method achieves state-of-the-art in both clean and various noisy conditions. Furthermore, we demonstrate the better generality of MSRL system than other baselines when test set contains unseen noises.
翻译:视听语音识别(AVSR)因能改善语音识别的噪声鲁棒性而取得了显著成功。主流方法侧重于融合音频与视觉输入以获取模态不变表征。然而,此类表征易过度依赖音频模态——在纯净条件下,音频模态比视频模态更易识别,导致AVSR模型在面对噪声干扰时低估视觉流的重要性。为此,我们利用视觉模态特定表征为AVSR任务提供稳定的互补信息。具体而言,我们提出一种基于强化学习(RL)的框架MSRL,其中智能体在自回归解码过程中动态协调模态不变表征与模态特定表征。我们定制了与任务特定指标(即词错误率)直接相关的奖励函数,促使MSRL有效探索最优融合策略。在LRS3数据集上的实验结果表明,所提方法在纯净及多种噪声条件下均达到最优水平。此外,我们证明了当测试集包含未见噪声时,MSRL系统相比其他基线具有更好的泛化性。