As a pivotal task in natural language processing, element extraction has gained significance in the legal domain. Extracting legal elements from judicial documents helps enhance interpretative and analytical capacities of legal cases, and thereby facilitating a wide array of downstream applications in various domains of law. Yet existing element extraction datasets are limited by their restricted access to legal knowledge and insufficient coverage of labels. To address this shortfall, we introduce a more comprehensive, large-scale criminal element extraction dataset, comprising 15,831 judicial documents and 159 labels. This dataset was constructed through two main steps: first, designing the label system by our team of legal experts based on prior legal research which identified critical factors driving and processes generating sentencing outcomes in criminal cases; second, employing the legal knowledge to annotate judicial documents according to the label system and annotation guideline. The Legal Element ExtraCtion dataset (LEEC) represents the most extensive and domain-specific legal element extraction dataset for the Chinese legal system. Leveraging the annotated data, we employed various SOTA models that validates the applicability of LEEC for Document Event Extraction (DEE) task. The LEEC dataset is available on https://github.com/THUlawtech/LEEC .
翻译:作为自然语言处理中的关键任务,要素抽取在法律领域的重要性日益凸显。从司法文书中提取法律要素有助于增强法律案件的可解释性与分析能力,进而促进法律各领域下游应用的广泛发展。然而,现有要素抽取数据集受限于其有限的法律知识获取途径以及标签覆盖范围不足。为弥补这一缺陷,我们提出一个更全面、大规模的刑事要素抽取数据集,包含15,831份司法文书和159个标签。该数据集通过两个主要步骤构建:首先,由我们的法律专家团队基于先期法律研究设计标签体系,该研究识别了刑事案件中驱动量刑结果的关键因素与生成过程;其次,依据标签体系与标注指南,运用法律知识对司法文书进行标注。法律要素抽取数据集(LEEC)是针对中国法律体系最全面且具有领域专用性的法律要素抽取数据集。利用已标注数据,我们采用多种SOTA模型验证了LEEC在文档事件抽取(DEE)任务中的适用性。LEEC数据集可通过https://github.com/THUlawtech/LEEC 获取。