Recent work has shown that, while large language models (LLMs) demonstrate strong word translation or bilingual lexicon induction (BLI) capabilities in few-shot setups, they still cannot match the performance of 'traditional' mapping-based approaches in the unsupervised scenario where no seed translation pairs are available, especially for lower-resource languages. To address this challenge with LLMs, we propose self-augmented in-context learning (SAIL) for unsupervised BLI: starting from a zero-shot prompt, SAIL iteratively induces a set of high-confidence word translation pairs for in-context learning (ICL) from an LLM, which it then reapplies to the same LLM in the ICL fashion. Our method shows substantial gains over zero-shot prompting of LLMs on two established BLI benchmarks spanning a wide range of language pairs, also outperforming mapping-based baselines across the board. In addition to achieving state-of-the-art unsupervised BLI performance, we also conduct comprehensive analyses on SAIL and discuss its limitations.
翻译:最近研究表明,尽管大语言模型在少样本设置下展现出强大的词汇翻译或双语词汇归纳能力,但在无种子翻译对的无监督场景中——尤其对于低资源语言——其性能仍不及基于映射的传统方法。为应对大语言模型面临的这一挑战,我们提出面向无监督双语词汇归纳的自增强上下文学习(SAIL):从零样本提示出发,SAIL通过迭代方式从大语言模型中诱导生成高置信度的词汇翻译对用于上下文学习,再以上下文学习范式将这些翻译对重新应用于同一大语言模型。在两个覆盖广泛语言对的权威双语词汇归纳基准测试中,我们的方法较之大语言模型的零样本提示取得显著提升,并全面超越基于映射的基线方法。除实现无监督双语词汇归纳的最新性能外,我们还对SAIL进行了全面分析并讨论了其局限性。