Bilingual word lexicons are crucial tools for multilingual natural language understanding and machine translation tasks, as they facilitate the mapping of words in one language to their synonyms in another language. To achieve this, numerous papers have explored bilingual lexicon induction (BLI) in high-resource scenarios, using a typical pipeline consisting of two unsupervised steps: bitext mining and word alignment, both of which rely on pre-trained large language models~(LLMs). In this paper, we present an analysis of the BLI pipeline for German and two of its dialects, Bavarian and Alemannic. This setup poses several unique challenges, including the scarcity of resources, the relatedness of the languages, and the lack of standardization in the orthography of dialects. To evaluate the BLI outputs, we analyze them with respect to word frequency and pairwise edit distance. Additionally, we release two evaluation datasets comprising 1,500 bilingual sentence pairs and 1,000 bilingual word pairs. They were manually judged for their semantic similarity for each Bavarian-German and Alemannic-German language pair.
翻译:双语词汇表是多语言自然语言理解和机器翻译任务中的关键工具,它们能够将一种语言的词汇映射到另一种语言的同义词上。为此,大量研究在高资源场景下探索了双语词汇表归纳(BLI),通常采用由两个无监督步骤组成的标准流程:双语文本挖掘和词对齐,这两个步骤均依赖于预训练的大语言模型(LLM)。本文针对德语及其两种方言——巴伐利亚语和阿勒曼尼语——对BLI流程进行了分析。这一设置带来了若干独特挑战,包括资源匮乏、语言之间的亲属关系以及方言正字法的缺乏标准化。为了评估BLI的输出,我们从词频和成对编辑距离角度对其进行了分析。此外,我们发布了两个评估数据集,分别包含1,500个双语句子对和1,000个双语单词对,这些数据由人工对每个巴伐利亚语-德语和阿勒曼尼语-德语语言对的语义相似度进行了标注。