Large Language Models (LLMs) such as GPT-3 have emerged as general-purpose language models capable of addressing many natural language generation or understanding tasks. On the task of Machine Translation (MT), multiple works have investigated few-shot prompting mechanisms to elicit better translations from LLMs. However, there has been relatively little investigation on how such translations differ qualitatively from the translations generated by standard Neural Machine Translation (NMT) models. In this work, we investigate these differences in terms of the literalness of translations produced by the two systems. Using literalness measures involving word alignment and monotonicity, we find that translations out of English (E-X) from GPTs tend to be less literal, while exhibiting similar or better scores on MT quality metrics. We demonstrate that this finding is borne out in human evaluations as well. We then show that these differences are especially pronounced when translating sentences that contain idiomatic expressions.
翻译:大型语言模型(LLM)如GPT-3已涌现为通用语言模型,能够处理众多自然语言生成或理解任务。在机器翻译任务上,多项研究探索了少样本提示机制以引导LLM生成更优质的译文。然而,关于这类译文与标准神经机器翻译模型生成的译文在质性上存在何种差异,相关研究仍相对匮乏。本研究从两种系统所生成译文的直译程度入手,考察这些差异。通过采用涉及词汇对齐和平直度的直译度量方式,我们发现GPT模型从英语向外翻译时倾向于产生较少直译的译文,同时在机器翻译质量指标上表现相近或更优。我们通过人工评估进一步验证了这一发现。研究表明,当翻译含有习语表达的句子时,这种差异尤为显著。