The development of large-scale Chinese language models is flourishing, yet there is a lack of corresponding capability assessments. Therefore, we propose a test to measure the multitask accuracy of large Chinese language models. This test encompasses four major domains, including medicine, law, psychology, and education, with 15 subtasks in medicine and 8 subtasks in education. We found that the best-performing models in the zero-shot setting outperformed the worst-performing models by nearly 18.6 percentage points on average. Across the four major domains, the highest average zero-shot accuracy of all models is 0.512. In the subdomains, only the GPT-3.5-turbo model achieved a zero-shot accuracy of 0.693 in clinical medicine, which was the highest accuracy among all models across all subtasks. All models performed poorly in the legal domain, with the highest zero-shot accuracy reaching only 0.239. By comprehensively evaluating the breadth and depth of knowledge across multiple disciplines, this test can more accurately identify the shortcomings of the models.
翻译:大规模中文语言模型的发展方兴未艾,但相应的能力评估却存在缺失。为此,我们提出一项测试,用于衡量大型中文语言模型的多任务准确率。该测试涵盖医学、法律、心理学和教育四大领域,其中医学包含15个子任务,教育包含8个子任务。研究发现,在零样本场景下,表现最佳的模型平均准确率比最差模型高出近18.6个百分点。在四大领域中,所有模型的最高平均零样本准确率为0.512。在细分领域,仅GPT-3.5-turbo模型在临床医学领域达到0.693的零样本准确率,这也是所有模型在所有子任务中的最高准确率。所有模型在法律领域表现均不理想,最高零样本准确率仅为0.239。通过全面评估多个学科的知识广度与深度,本测试能够更精准地识别模型的不足之处。