This paper tackles the task of legal extractive summarization using a dataset of 430K U.S. court opinions with key passages annotated. According to automated summary quality metrics, the reinforcement-learning-based MemSum model is best and even out-performs transformer-based models. In turn, expert human evaluation shows that MemSum summaries effectively capture the key points of lengthy court opinions. Motivated by these results, we open-source our models to the general public. This represents progress towards democratizing law and making U.S. court opinions more accessible to the general public.
翻译:本文利用包含43万份美国法院判决书及关键段落标注的数据集,研究法律文本的抽取式摘要任务。基于自动摘要质量评估指标,采用强化学习的MemSum模型表现最佳,甚至优于基于Transformer的模型。专家人工评估表明,MemSum生成的摘要能有效捕捉长篇判决书的核心要点。基于上述成果,我们向公众开源了相关模型。这标志着在推进法律民主化、提升美国法院判决书公众可及性方面取得了重要进展。