Large language models have recently made tremendous progress in a variety of aspects, e.g., cross-task generalization, instruction following. Comprehensively evaluating the capability of large language models in multiple tasks is of great importance. In this paper, we propose M3KE, a Massive Multi-Level Multi-Subject Knowledge Evaluation benchmark, which is developed to measure knowledge acquired by Chinese large language models by testing their multitask accuracy in zero- and few-shot settings. We have collected 20,477 questions from 71 tasks. Our selection covers all major levels of Chinese education system, ranging from the primary school to college, as well as a wide variety of subjects, including humanities, history, politics, law, education, psychology, science, technology, art and religion. All questions are multiple-choice questions with four options, hence guaranteeing a standardized and unified assessment process. We've assessed a number of state-of-the-art open-source Chinese large language models on the proposed benchmark. The size of these models varies from 335M to 130B parameters. Experiment results demonstrate that they perform significantly worse than GPT-3.5 that reaches an accuracy of ~ 48% on M3KE. The dataset is available at https://github.com/tjunlp-lab/M3KE.
翻译:大型语言模型近期在跨任务泛化、指令遵循等多个方面取得了巨大进展。全面评估大型语言模型在多任务中的能力具有重要意义。本文提出M3KE(大规模多层多学科知识评估基准),旨在通过测试中文大语言模型在零样本和少样本设置下的多任务准确率,衡量其获取的知识水平。我们从71个任务中收集了20,477道题目,涵盖中国教育体系从小学到大学的所有主要阶段,以及人文、历史、政治、法律、教育、心理学、科学、技术、艺术和宗教等多个学科领域。所有题目均为四选一选择题,从而保证标准化且统一的评估流程。我们基于该基准评估了多个开源中文大语言模型,模型参数量从3.35亿到1300亿不等。实验结果表明,这些模型的性能显著低于GPT-3.5(其在M3KE上的准确率约为48%)。数据集已公开于https://github.com/tjunlp-lab/M3KE。