Document-level Relation Extraction (DocRE), which aims to extract relations from a long context, is a critical challenge in achieving fine-grained structural comprehension and generating interpretable document representations. Inspired by recent advances in in-context learning capabilities emergent from large language models (LLMs), such as ChatGPT, we aim to design an automated annotation method for DocRE with minimum human effort. Unfortunately, vanilla in-context learning is infeasible for document-level relation extraction due to the plenty of predefined fine-grained relation types and the uncontrolled generations of LLMs. To tackle this issue, we propose a method integrating a large language model (LLM) and a natural language inference (NLI) module to generate relation triples, thereby augmenting document-level relation datasets. We demonstrate the effectiveness of our approach by introducing an enhanced dataset known as DocGNRE, which excels in re-annotating numerous long-tail relation types. We are confident that our method holds the potential for broader applications in domain-specific relation type definitions and offers tangible benefits in advancing generalized language semantic comprehension.
翻译:文档级关系抽取(DocRE)旨在从长上下文中提取关系,是实现细粒度结构理解与可解释文档表示的关键挑战。受近期大语言模型(如ChatGPT)在上下文学习能力方面取得进展的启发,我们致力于设计一种以最少人工干预实现DocRE自动化标注的方法。然而,由于预定义的细粒度关系类型繁多且大语言模型生成内容不可控,传统上下文学习难以直接应用于文档级关系抽取。为解决该问题,我们提出了一种融合大语言模型与自然语言推理模块的方法,用于生成关系三元组,从而增强文档级关系数据集。通过引入增强数据集DocGNRE,我们验证了该方法在重新标注大量长尾关系类型方面的有效性。我们相信,本方法在特定领域关系类型定义中具有更广泛的适用潜力,并为推动通用语言语义理解提供了切实优势。