Over the last few years, large language models (LLMs) have emerged as the most important breakthroughs in natural language processing (NLP) that fundamentally transform research and developments in the field. ChatGPT represents one of the most exciting LLM systems developed recently to showcase impressive skills for language generation and highly attract public attention. Among various exciting applications discovered for ChatGPT in English, the model can process and generate texts for multiple languages due to its multilingual training data. Given the broad adoption of ChatGPT for English in different problems and areas, a natural question is whether ChatGPT can also be applied effectively for other languages or it is necessary to develop more language-specific technologies. The answer to this question requires a thorough evaluation of ChatGPT over multiple tasks with diverse languages and large datasets (i.e., beyond reported anecdotes), which is still missing or limited in current research. Our work aims to fill this gap for the evaluation of ChatGPT and similar LLMs to provide more comprehensive information for multilingual NLP applications. While this work will be an ongoing effort to include additional experiments in the future, our current paper evaluates ChatGPT on 7 different tasks, covering 37 diverse languages with high, medium, low, and extremely low resources. We also focus on the zero-shot learning setting for ChatGPT to improve reproducibility and better simulate the interactions of general users. Compared to the performance of previous models, our extensive experimental results demonstrate a worse performance of ChatGPT for different NLP tasks and languages, calling for further research to develop better models and understanding for multilingual learning.
翻译:近年来,大语言模型(LLM)作为自然语言处理(NLP)领域最重要的突破性成果,从根本上改变了该领域的研究与发展格局。ChatGPT是近期开发的最具代表性的LLM系统之一,其展现出的卓越语言生成能力引发了公众的广泛关注。在已发现的英语多样化应用中,由于训练数据涵盖多语言,该模型可处理并生成多种语言的文本。鉴于ChatGPT在英语环境下不同领域和问题中的广泛应用,一个自然的问题随之产生:ChatGPT能否有效适用于其他语言,抑或仍需开发更多面向特定语言的技术?回答该问题需要对ChatGPT在多种任务、多语言及大规模数据集(即超越现有零散案例报道)中进行系统评估,而当前研究仍缺乏此类评估或存在局限性。本研究旨在填补ChatGPT及相关LLM评估的空白,为多语言NLP应用提供更全面的信息。尽管本项目将持续推进后续实验,当前论文已在涵盖高资源、中资源、低资源及极低资源共计37种语言的7项不同任务上对ChatGPT进行了评估。我们特别采用零样本学习设置以提升可复现性并模拟普通用户交互场景。与先前模型性能相比,广泛实验结果表明ChatGPT在不同NLP任务及语言中的表现欠佳,这呼吁学界开展进一步研究以构建更优的多语言学习模型与认知框架。