Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge benchmarks, but across many languages. We show such benchmarks, and consequently multilingual evaluations, measure mathematical reasoning and factual recall, not multilingual proficiency. For example, thinking variants dramatically outperform instruct variants on these benchmarks, yet often perform worse on real-world multilingual tasks, such as LMArena. We propose a simple alternative: evaluate multilingual capability via round-trip translation. Given text in a source language, translate it to a target language and back; semantic gaps between the original and result expose failures in multilingual generation capabilities. Round-trip translation correlates almost perfectly (\r{ho} = 0.94) with user ratings on LMArena with our benchmark, requires no human reference translations, and does not require a more capable multilingual judge than tested models. Lastly, we introduce Lost in Translation (LiT), a challenging round-trip translation benchmark spanning widely spoken languages worldwide, for realistic evaluation of multilingual frontier models.
翻译:多语言基准测试指导着前沿模型的发展。然而,前沿模型报告的多语言评估与流行的推理和知识基准测试结构相似,但覆盖多种语言。我们证明,此类基准测试及其相应的多语言评估衡量的是数学推理和事实回忆能力,而非多语言熟练程度。例如,推理变体在这些基准测试中显著优于指令变体,但在真实多语言任务(如LMArena)中往往表现更差。我们提出一个简单替代方案:通过往返翻译评估多语言能力。给定源语言文本,将其翻译为目标语言再译回源语言;原文与译文之间的语义差距揭示了多语言生成能力的不足。往返翻译与用户对LMArena评分的相关性近乎完美(ρ=0.94),无需人工参考翻译,也无需比测试模型更强大的多语言评估器。最后,我们引入"迷失在翻译中"(LiT)——一个覆盖全球广泛使用语言的具有挑战性的往返翻译基准测试,用于对多语言前沿模型进行现实评估。