With the emergence of numerous legal LLMs, there is currently a lack of a comprehensive benchmark for evaluating their legal abilities. In this paper, we propose the first Chinese Legal LLMs benchmark based on legal capabilities. Through the collaborative efforts of legal and artificial intelligence experts, we divide the legal capabilities of LLMs into three levels: basic legal NLP capability, basic legal application capability, and complex legal application capability. We have completed the first phase of evaluation, which mainly focuses on the capability of basic legal NLP. The evaluation results show that although some legal LLMs have better performance than their backbones, there is still a gap compared to ChatGPT. Our benchmark can be found at URL.
翻译:随着大量法律大语言模型的出现,目前缺乏一个用于评估其法律能力的全面基准。本文提出了首个基于法律能力的中文法律大语言模型基准。通过法律和人工智能专家的协同努力,我们将大语言模型的法律能力划分为三个层次:基础法律自然语言处理能力、基础法律应用能力和复杂法律应用能力。我们完成了第一阶段的评估,主要聚焦于基础法律自然语言处理能力。评估结果表明,尽管部分法律大语言模型相比其骨干模型表现更优,但与ChatGPT相比仍存在差距。我们的基准可在URL获取。