Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there are thousands of spoken languages worldwide, and adapting to new languages is an important problem. In this work, we aim to understand which model adapts better to languages unseen during pre-training. We fine-tune both models on 13 unseen languages and 18 seen languages. Our results show that the number of hours seen per language and language family during pre-training is predictive of how the models compare, despite the significant differences in the pre-training methods.
翻译:近年来,XLS-R和Whisper等模型通过分别在大约100种口语语言的音频上进行预训练,显著提升了多语言语音技术的可及性。然而,全球存在数千种口语语言,如何将模型适应新语言仍是一个重要课题。本研究旨在探究哪种模型能更好地适应预训练中未见的语言。我们对两种模型在13种未见语言和18种已见语言上进行了微调实验。结果表明,尽管预训练方法存在显著差异,但预训练过程中每种语言及其语系所对应的音频小时数,能够有效预测模型在不同语言上的性能比较结果。