Large language models (LLMs) are competitive with the state of the art on a wide range of sentence-level translation datasets. However, their ability to translate paragraphs and documents remains unexplored because evaluation in these settings is costly and difficult. We show through a rigorous human evaluation that asking the Gpt-3.5 (text-davinci-003) LLM to translate an entire literary paragraph (e.g., from a novel) at once results in higher-quality translations than standard sentence-by-sentence translation across 18 linguistically-diverse language pairs (e.g., translating into and out of Japanese, Polish, and English). Our evaluation, which took approximately 350 hours of effort for annotation and analysis, is conducted by hiring translators fluent in both the source and target language and asking them to provide both span-level error annotations as well as preference judgments of which system's translations are better. We observe that discourse-level LLM translators commit fewer mistranslations, grammar errors, and stylistic inconsistencies than sentence-level approaches. With that said, critical errors still abound, including occasional content omissions, and a human translator's intervention remains necessary to ensure that the author's voice remains intact. We publicly release our dataset and error annotations to spur future research on evaluation of document-level literary translation.
翻译:大型语言模型(LLMs)在广泛的句子级翻译数据集上已达到与当前最优方法相竞争的水平。然而,由于段落和文档级翻译的评估成本高昂且难度较大,其在这类场景下的翻译能力尚未得到充分探索。我们通过一项严谨的人工评估表明,在18个语言多样性丰富的语对中(例如涉及日语、波兰语和英语之间的双向翻译),要求Gpt-3.5(text-davinci-003)等大型语言模型一次性翻译整个文学段落(如小说选段),其翻译质量显著优于逐句翻译方法。本次评估耗时约350小时用于标注与分析,我们聘请了精通源语言和目标语言的译者,要求他们提供跨层级错误标注以及译文质量偏好判断。研究发现,文档级LLM翻译器在减少误译、语法错误和文体不一致性方面显著优于句子级方法。尽管如此,关键错误仍然频繁出现,包括偶发的内容遗漏,且人类译者的介入对于保持作者原声的完整性仍是必要的。为促进文档级文学翻译评估的未来研究,我们公开了相关数据集及错误标注。