Developing effective spoken language processing systems for low-resource languages poses several challenges due to the lack of parallel data and limited resources for fine-tuning models. In this work, we target on improving upon both text classification and translation of Nigerian Pidgin (Naija) by collecting a large-scale parallel English-Pidgin corpus and further propose a framework of cross-lingual adaptive training that includes both continual and task adaptive training so as to adapt a base pre-trained model to low-resource languages. Our studies show that English pre-trained language models serve as a stronger prior than multilingual language models on English-Pidgin tasks with up to 2.38 BLEU improvements; and demonstrate that augmenting orthographic data and using task adaptive training with back-translation can have a significant impact on model performance.
翻译:针对低资源语言开发高效的口语处理系统面临诸多挑战,主要源于并行数据匮乏及模型微调资源有限。本文通过构建大规模英-皮平行语料库,致力于提升尼日利亚皮钦语(Naija)的文本分类与翻译性能,并提出包含持续自适应训练与任务自适应训练的跨语言自适应训练框架,以将基础预训练模型适配至低资源语言。研究表明,在英-皮任务中,英语预训练语言模型比多语言语言模型具有更强的先验知识,BLEU值提升最高达2.38;同时证明,增强正字法数据并采用基于反向翻译的任务自适应训练可显著影响模型性能。