We investigate how large language models perform on low-resource languages by benchmarking eight LLMs across five experimental conditions in English, Kazakh, and Mongolian. Using 50 hand-crafted questions spanning factual, reasoning, technical, and culturally grounded categories, we evaluate 2,000 responses on accuracy, fluency, and completeness. We find a consistent performance gap of 13.8-16.7 percentage points between English and low-resource language conditions, with models maintaining surface-level fluency while producing significantly less accurate content. Cross-lingual transfer-prompting models to reason in English before translating back-yields selective gains for bilingual architectures (+2.2pp to +4.3pp) but provides no benefit to English-dominant models. Our results demonstrate that current LLMs systematically underserve low-resource language communities, and that effective mitigation strategies are architecture-dependent rather than universal.
翻译:我们通过基准测试八个大型语言模型在英语、哈萨克语和蒙古语五种实验条件下的表现,研究了大型语言模型处理低资源语言的能力。利用50个手工设计的问题,涵盖事实性、推理性、技术性和文化基础性类别,我们对2000个回答在准确性、流畅性和完整性方面进行了评估。我们发现英语与低资源语言条件之间存在13.8至16.7个百分点的持续性能差距,模型在保持表面流畅性的同时,生成的内容准确性显著降低。跨语言迁移——提示模型先用英语推理再回译——对双语架构模型带来选择性提升(+2.2至+4.3个百分点),但对以英语为主的模型没有益处。我们的结果表明,当前大型语言模型系统性地未能充分服务于低资源语言社区,且有效的缓解策略依赖于架构而非具有普适性。