The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both languages. However, in reality, there can be an imbalance among the languages for the available spoken captions. Our key contribution in this work is to leverage the power of a high-resource language in a bilingual visually grounded speech model to improve the performance of a low-resource language. We introduce two methods to distill the knowledge of high-resource language into low-resource languages: (1) incorporating a strong pre-trained high-resource language encoder and (2) using semantically similar spoken captions. Our experiments show that combining these two approaches effectively enables the low-resource language to surpass the performances of monolingual and bilingual counterparts for cross-modal retrieval tasks.
翻译:本研究的目的是从多语言视角探索视觉关联语音模型(VGS)的学习过程。双语VGS模型通常使用两种语言数量相等的语音标注进行训练,然而现实情况下,不同语言可获得的语音标注数量可能存在不平衡。本文的核心贡献在于利用双语视觉关联语音模型中高资源语言的能力,提升低资源语言的性能。我们提出了两种将高资源语言知识迁移到低资源语言的方法:(1)引入强大的预训练高资源语言编码器;(2)使用语义相似的语音标注。实验表明,将这两种方法结合使用,能有效使低资源语言在跨模态检索任务中超越单语模型和双语模型的性能。