The growth of pending legal cases in populous countries, such as India, has become a major issue. Developing effective techniques to process and understand legal documents is extremely useful in resolving this problem. In this paper, we present our systems for SemEval-2023 Task 6: understanding legal texts (Modi et al., 2023). Specifically, we first develop the Legal-BERT-HSLN model that considers the comprehensive context information in both intra- and inter-sentence levels to predict rhetorical roles (subtask A) and then train a Legal-LUKE model, which is legal-contextualized and entity-aware, to recognize legal entities (subtask B). Our evaluations demonstrate that our designed models are more accurate than baselines, e.g., with an up to 15.0% better F1 score in subtask B. We achieved notable performance in the task leaderboard, e.g., 0.834 micro F1 score, and ranked No.5 out of 27 teams in subtask A.
翻译:在人口众多的国家(如印度),待审案件数量的增长已成为一个重大问题。开发有效的法律文件处理与理解技术对于解决这一难题极为重要。本文介绍了我们针对SemEval-2023任务6(理解法律文本,Modi等人,2023)所构建的系统。具体而言,我们首先开发了Legal-BERT-HSLN模型,该模型在句内和句间层面综合考虑完整的上下文信息以预测修辞角色(子任务A),随后训练了具备法律情境化与实体感知能力的Legal-LUKE模型,用于识别法律实体(子任务B)。评估结果表明,我们设计的模型相比基线方法具有更高准确性,例如在子任务B中F1分数提升高达15.0%。我们在任务排行榜上取得了显著成绩,例如子任务A中以0.834的微平均F1分数在27支参赛队伍中排名第5。