The escalating number of pending cases is a growing concern world-wide. Recent advancements in digitization have opened up possibilities for leveraging artificial intelligence (AI) tools in the processing of legal documents. Adopting a structured representation for legal documents, as opposed to a mere bag-of-words flat text representation, can significantly enhance processing capabilities. With the aim of achieving this objective, we put forward a set of diverse attributes for criminal case proceedings. We use a state-of-the-art sequence labeling framework to automatically extract attributes from the legal documents. Moreover, we demonstrate the efficacy of the extracted attributes in a downstream task, namely legal judgment prediction.
翻译:待审案件数量的持续上升是全球日益关注的焦点。近年来数字化技术的进步为利用人工智能工具处理法律文件提供了可能。采用结构化表示(而非简单的词袋扁平文本表示)来呈现法律文件,能够显著提升处理能力。为实现这一目标,我们针对刑事案件诉讼流程提出了一套多样化的属性集,并采用先进的序列标注框架从法律文件中自动提取这些属性。此外,我们通过下游任务——即法律判决预测——验证了所提取属性的有效性。