Translating literary works has perennially stood as an elusive dream in machine translation (MT), a journey steeped in intricate challenges. To foster progress in this domain, we hold a new shared task at WMT 2023, the first edition of the Discourse-Level Literary Translation. First, we (Tencent AI Lab and China Literature Ltd.) release a copyrighted and document-level Chinese-English web novel corpus. Furthermore, we put forth an industry-endorsed criteria to guide human evaluation process. This year, we totally received 14 submissions from 7 academia and industry teams. We employ both automatic and human evaluations to measure the performance of the submitted systems. The official ranking of the systems is based on the overall human judgments. In addition, our extensive analysis reveals a series of interesting findings on literary and discourse-aware MT. We release data, system outputs, and leaderboard at http://www2.statmt.org/wmt23/literary-translation-task.html.
翻译:文学作品的翻译一直是机器翻译中难以企及的梦想,这项任务充满了错综复杂的挑战。为促进该领域的发展,我们在WMT 2023上首次举办了篇章级文学翻译共享任务。首先,我们(腾讯人工智能实验室与中国文学有限公司)发布了一个受版权保护的文档级中英网络小说语料库。此外,我们还提出了一套业内认可的标准来指导人工评估过程。今年,我们共收到来自7个学术界和工业界团队的14份提交成果。我们采用自动评估与人工评估相结合的方法来评测各提交系统的性能,并根据总体人工评判结果确定系统排名。进一步的深入分析揭示了文学翻译与篇章感知机器翻译领域的一系列有趣发现。我们在http://www2.statmt.org/wmt23/literary-translation-task.html上发布数据集、系统输出结果及排行榜。