The release of top-performing open-weight LLMs has cemented China's role as a leading force in AI development. Do these models support languages spoken in China? Or do they support the same languages as models developed in the United States or in Europe? Comparing multilingual capabilities is important for two reasons. First, language ability provides insights into pre-training data curation, and thus into resource allocation and development priorities. Second, Chinese model developers need to navigate the tension between serving a linguistically diverse population domestically, and optimizing for globally visible benchmarks that are predominantly English. We investigate Chinese model developers' priorities through a comparative study of Chinese-developed and Western-developed open-weight LLMs, on 21 language variants including Asian regional, Chinese, and European languages. Our experiments on Information Parity and reading comprehension show Chinese models' performance across these languages correlates strongly (r=0.93) with their Western counterparts, with the sole exception being better Mandarin. Chinese-developed models are good at French and German, but they sometimes cannot identify languages spoken by Chinese minorities such as Kazakh and Uyghur. Overall, all open-weight LLMs we study have a similar multilingual performance profile, despite the diverse linguistic and cultural contexts the model developers operated within. We interpret the homogenization as consistent with the influence of global benchmarking practices and shared training resources. Rather than treating current language support as inevitable, our results highlight multilingual development as a space of prioritization and trade-offs, with implications for model developers, policymakers, and users.
翻译:顶级性能开源权重LLM的发布巩固了中国作为AI发展领导者的地位。这些模型是否支持中国境内使用的语言?还是支持与美国或欧洲开发的模型相同的语言?比较多语言能力至关重要,原因有二。首先,语言能力揭示了预训练数据整理方式,进而反映资源分配与发展重点。其次,中国模型开发者需平衡国内语言多样性需求与主要面向英语的全球基准优化。我们通过对比中西方开发的开源权重LLM,在涵盖亚洲区域语言、中国语言及欧洲语言的21种语言变体上,探究中国模型开发者的优先级。在信息对等性和阅读理解实验中发现,中国模型在这些语言上的表现与西方模型高度相关(r=0.93),唯一例外是普通话表现更优。中国开发的模型擅长法语和德语,但有时无法识别哈萨克语、维吾尔语等中国少数民族语言。总体而言,尽管开发者身处不同的语言文化背景,所有研究的开源权重LLM均呈现相似的多语言能力分布。我们将这种同质化现象视为全球基准实践与共享训练资源共同作用的结果。我们不应将当前语言支持视为必然结果,而应视多语言发展为优先级选择与权衡取舍的空间——这对模型开发者、政策制定者及用户均具有深远意义。