Due to the limited availability of high quality datasets for training sentence embeddings in Turkish, we propose a training methodology and a regimen to develop a sentence embedding model. The central idea is simple but effective : is to fine-tune a pretrained encoder-decoder model in two consecutive stages, where the first stage involves aligning the embedding space with translation pairs. Thanks to this alignment, the prowess of the main model can be better projected onto the target language in a sentence embedding setting where it can be fine-tuned with high accuracy in short duration with limited target language dataset.
翻译:由于用于训练土耳其语句子嵌入的高质量数据集有限,我们提出了一种训练方法和策略来开发句子嵌入模型。核心思想简单而有效:即分两个连续阶段微调预训练的编码器-解码器模型,其中第一阶段涉及将嵌入空间与翻译对进行对齐。借助这种对齐,主模型的能力可以更好地投射到目标语言上,使得在仅使用有限的目标语言数据集时,能够在短时间内以高精度进行微调,从而应用于句子嵌入场景。