Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under diverse conditions even on languages it was evaluated on such as Arabic. In this work, we address this gap by comprehensively evaluating Whisper on several varieties of Arabic speech for the ASR task. Our evaluation covers most publicly available Arabic speech data and is performed under n-shot (zero-, few-, and full) finetuning. We also investigate the robustness of Whisper under completely novel conditions, such as in dialect-accented standard Arabic and in unseen dialects for which we develop evaluation data. Our experiments show that although Whisper zero-shot outperforms fully finetuned XLS-R models on all datasets, its performance deteriorates significantly in the zero-shot setting for five unseen dialects (i.e., Algeria, Jordan, Palestine, UAE, and Yemen).
翻译:Whisper作为近期开发的多语言弱监督模型,据报告在单语和多语场景下的多项语音识别基准测试中表现优异。然而,即便在已评估的语言(如阿拉伯语)上,Whisper在多样化条件下的表现尚不明确。本研究通过全面评估Whisper在多种阿拉伯语变体上的自动语音识别(ASR)任务,填补了这一研究空白。我们的评估覆盖了大部分公开可用的阿拉伯语语音数据,并在零样本、少样本和全样本三种微调范式下进行。同时,我们探讨了Whisper在全新条件下的鲁棒性,包括带方言口音的标准阿拉伯语及未见方言(为此我们构建了评估数据集)。实验表明:尽管Whisper的零样本性能在所有数据集上均超越完全微调的XLS-R模型,但在阿尔及利亚、约旦、巴勒斯坦、阿联酋和也门这五种未见方言的零样本场景下,其性能显著下降。