Contrastive learning has been successfully used for retrieval of semantically aligned sentences, but it often requires large batch sizes or careful engineering to work well. In this paper, we instead propose a generative model for learning multilingual text embeddings which can be used to retrieve or score sentence pairs. Our model operates on parallel data in $N$ languages and, through an approximation we introduce, efficiently encourages source separation in this multilingual setting, separating semantic information that is shared between translations from stylistic or language-specific variation. We show careful large-scale comparisons between contrastive and generation-based approaches for learning multilingual text embeddings, a comparison that has not been done to the best of our knowledge despite the popularity of these approaches. We evaluate this method on a suite of tasks including semantic similarity, bitext mining, and cross-lingual question retrieval -- the last of which we introduce in this paper. Overall, our Variational Multilingual Source-Separation Transformer (VMSST) model outperforms both a strong contrastive and generative baseline on these tasks.
翻译:对比学习已成功用于检索语义对齐的句子,但往往需要大批量或精细的工程才能良好运行。本文提出了一种生成模型,用于学习多语言文本嵌入,可用于检索或评分句子对。我们的模型在$N$种语言的平行数据上运行,并通过我们引入的近似方法,在多语言环境中有效促进源分离,将翻译间共享的语义信息与风格或语言特定变异区分开来。我们进行了对比学习和基于生成的方法之间的细致大规模比较,用于学习多语言文本嵌入,尽管这些方法已流行,但据我们所知,这种比较尚未进行。我们在包括语义相似性、双语文本挖掘和跨语言问题检索(本文首次引入)等一系列任务上评估了该方法。总体而言,我们的变分多语言源分离变换器(VMSST)模型在这些任务上均优于强对比学习和生成基线。