AV-HuBERT, a multi-modal self-supervised learning model, has been shown to be effective for categorical problems such as automatic speech recognition and lip-reading. This suggests that useful audio-visual speech representations can be obtained via utilizing multi-modal self-supervised embeddings. Nevertheless, it is unclear if such representations can be generalized to solve real-world multi-modal AV regression tasks, such as audio-visual speech enhancement (AVSE) and audio-visual speech separation (AVSS). In this study, we leveraged the pre-trained AV-HuBERT model followed by an SE module for AVSE and AVSS. Comparative experimental results demonstrate that our proposed model performs better than the state-of-the-art AVSE and traditional audio-only SE models. In summary, our results confirm the effectiveness of our proposed model for the AVSS task with proper fine-tuning strategies, demonstrating that multi-modal self-supervised embeddings obtained from AV-HuBERT can be generalized to audio-visual regression tasks.
翻译:AV-HuBERT作为一种多模态自监督学习模型,已被证明在自动语音识别和唇读等分类任务中具有有效性。这表明通过多模态自监督嵌入可以获得有用的音频-视觉语音表征。然而,此类表征能否泛化应用于音频-视觉语音增强(AVSE)和音频-视觉语音分离(AVSS)等真实世界多模态回归任务尚不明确。本研究在预训练AV-HuBERT模型后接语音增强模块(SE),分别用于AVSE与AVSS任务。对比实验结果表明,所提模型性能优于当前最先进的AVSE模型及传统纯音频语音增强模型。总体而言,本研究通过适当的微调策略验证了所提模型在AVSS任务中的有效性,证实从AV-HuBERT获取的多模态自监督嵌入可泛化应用于音频-视觉回归任务。