Self-Supervised Learning (SSL) models have demonstrated exceptional performance in various speech tasks, particularly in low-resource and multilingual domains. Recent works show that fusing diverse SSL models could achieve superior performance compared to using one SSL model. However, fusing models increases the overall parameter size, leading to higher computational costs. We propose EFFUSE, a novel approach that uses a single SSL model to mimic the features of multiple SSL models via prediction, resulting in a lightweight framework with competitive performance. Our experiments show that EFFUSE outperforms individual SSL models in multilingual speech recognition tasks. Our best performing model achieves an average SUPERB score increase of 63.5 (6.3%) from the SSL baselines in Multilingual Speech Universal PERformance Benchmark (ML-SUPERB), while decreasing parameter size on average by 317M parameters (49%) from the fusion models.
翻译:自监督学习模型已在多种语音任务中展现出卓越性能,尤其在低资源与多语言领域。近期研究表明,融合多种自监督学习模型相较于使用单一模型可获得更优性能。然而,模型融合会增加整体参数量,导致计算成本上升。本文提出EFFUSE,一种创新方法,通过预测机制使单个自监督学习模型模拟多个自监督学习模型的特征,从而构建具有竞争力的轻量化框架。实验表明,EFFUSE在多语言语音识别任务中优于单一自监督学习模型。在ML-SUPERB多语言语音通用性能基准测试中,我们性能最优的模型相较于自监督学习基线平均SUPERB分数提升63.5(6.3%),同时相较于融合模型平均减少3.17亿参数(49%)。