Large-scale language-agnostic sentence embedding models such as LaBSE (Feng et al., 2022) obtain state-of-the-art performance for parallel sentence alignment. However, these large-scale models can suffer from inference speed and computation overhead. This study systematically explores learning language-agnostic sentence embeddings with lightweight models. We demonstrate that a thin-deep encoder can construct robust low-dimensional sentence embeddings for 109 languages. With our proposed distillation methods, we achieve further improvements by incorporating knowledge from a teacher model. Empirical results on Tatoeba, United Nations, and BUCC show the effectiveness of our lightweight models. We release our lightweight language-agnostic sentence embedding models LEALLA on TensorFlow Hub.
翻译:大规模语言无关句子嵌入模型(如LaBSE,Feng等,2022)在平行句子对齐任务中取得了最先进的性能。然而,这类大规模模型可能存在推理速度和计算开销方面的不足。本研究系统探索了使用轻量级模型学习语言无关句子嵌入的方法。我们证明,薄深编码器能够为109种语言构建鲁棒的低维句子嵌入。通过提出的蒸馏方法,我们进一步利用教师模型的知识实现了性能提升。在Tatoeba、联合国和BUCC数据集上的实验结果表明了轻量级模型的有效性。我们在TensorFlow Hub上发布了轻量级语言无关句子嵌入模型LEALLA。