Large Language Models (LLMs) such as GPT-3 have emerged as general-purpose language models capable of addressing many natural language generation or understanding tasks. On the task of Machine Translation (MT), multiple works have investigated few-shot prompting mechanisms to elicit better translations from LLMs. However, there has been relatively little investigation on how such translations differ qualitatively from the translations generated by standard Neural Machine Translation (NMT) models. In this work, we investigate these differences in terms of the literalness of translations produced by the two systems. Using literalness measures involving word alignment and monotonicity, we find that translations out of English (E-X) from GPTs tend to be less literal, while exhibiting similar or better scores on MT quality metrics. We demonstrate that this finding is borne out in human evaluations as well. We then show that these differences are especially pronounced when translating sentences that contain idiomatic expressions.
翻译:大型语言模型(LLMs),如GPT-3,已作为通用语言模型出现,能够处理许多自然语言生成或理解任务。在机器翻译(MT)任务中,多项研究探讨了少样本提示机制,以从LLMs中引出更好的翻译。然而,关于这些翻译与标准神经机器翻译(NMT)模型生成的翻译在质量上有何差异,相关研究相对较少。在本工作中,我们从两个系统生成翻译的直译程度角度调查这些差异。通过使用涉及词对齐和单调性的直译度量,我们发现GPT从英语到其他语言(E-X)的翻译往往更不直译,同时在MT质量指标上表现相似或更优。我们证明,这一发现在人工评估中也得以体现。随后,我们展示这些差异在翻译包含习语表达的句子时尤为显著。