Pre-trained Transformer-based speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem of domain mismatch remains a challenge in this area, where the domain of the pre-training data might differ from that of the downstream labeled data used for fine-tuning. In multilingual tasks such as SLID, the pre-trained speech model may not support all the languages in the downstream task. To address this challenge, we propose self-supervised adaptive pre-training (SAPT) to adapt the pre-trained model to the target domain and languages of the downstream task. We apply SAPT to the XLSR-128 model and investigate the effectiveness of this approach for the SLID task. First, we demonstrate that SAPT improves XLSR performance on the FLEURS benchmark with substantial gains up to 40.1% for under-represented languages. Second, we apply SAPT on four different datasets in a few-shot learning setting, showing that our approach improves the sample efficiency of XLSR during fine-tuning. Our experiments provide strong empirical evidence that continual adaptation via self-supervision improves downstream performance for multilingual speech models.
翻译:基于Transformer的预训练语音模型在微调至自动语音识别和口语语言识别(SLID)等下游任务时表现出显著性能。然而,该领域仍面临领域不匹配的挑战,即预训练数据的领域可能与微调时使用的下游标注数据存在差异。在多语言任务(如SLID)中,预训练语音模型可能无法覆盖下游任务所需的所有语言。为解决此问题,我们提出自监督自适应预训练(SAPT),使预训练模型适配下游任务的目标领域和语言。我们将SAPT应用于XLSR-128模型,并探究该方法对SLID任务的有效性。首先,我们证明SAPT在FLEURS基准上提升了XLSR的性能,尤其是对低资源语言的提升高达40.1%。其次,我们在四个不同数据集的少样本学习场景中应用SAPT,表明该方法提升了XLSR在微调过程中的样本效率。实验提供了强有力的实证证据:通过自监督进行持续自适应能够提升多语言语音模型的下游任务性能。