We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA). The benchmark consists of 25 multiple-choice questions that require reasoning in order to be answered correctly. Each question is provided in official human translations to 43 languages and complemented with machine-translated counterparts (i.e., 2,150 data points in total). We evaluate two mainstream proprietary LLMs across languages, reasoning effort levels, and translation types in terms of their ability to answer the questions correctly. Our results show that modern LLMs can reason effectively across all evaluated languages, achieve accuracy comparable to human test-takers, with some performance variations across covered languages. We further find that machine-translated questions do not degrade accuracy relative to official human translations which suggests that high-quality machine translation (synthetic data) might often be adequate for large-scale multilingual reasoning evaluations where official translations are not available. Finally, we analyze token usage and related inference cost and find that LLMs usage in some languages is simultaneously more expensive and less accurate.
翻译:我们提出了mmPISA-bench,一个从经合组织国际学生评估项目(PISA)中提取的紧凑型高质量多语言推理基准。该基准包含25道需要正确推理才能作答的多项选择题。每道题均提供经官方人工翻译的43种语言版本,并辅以机器翻译版本(总计2,150个数据点)。我们评估了两款主流闭源大语言模型在不同语言、推理难度级别及翻译类型下的答题准确性。结果显示,现代大语言模型在所有评估语言中均能有效进行推理,准确率与人类应试者相当,但不同语言间存在一定性能差异。进一步发现,机器翻译试题相对于官方人工翻译并未降低准确率,这表明当缺乏官方翻译时,高质量机器翻译(合成数据)通常足以支持大规模多语言推理评估。最后,我们分析了令牌使用及相关推理成本,发现部分语言中大语言模型的使用成本更高且准确率更低。