Information Extraction (IE) seeks to derive structured information from unstructured texts, often facing challenges in low-resource scenarios due to data scarcity and unseen classes. This paper presents a review of neural approaches to low-resource IE from \emph{traditional} and \emph{LLM-based} perspectives, systematically categorizing them into a fine-grained taxonomy. Then we conduct empirical study on LLM-based methods compared with previous state-of-the-art models, and discover that (1) well-tuned LMs are still predominant; (2) tuning open-resource LLMs and ICL with GPT family is promising in general; (3) the optimal LLM-based technical solution for low-resource IE can be task-dependent. In addition, we discuss low-resource IE with LLMs, highlight promising applications, and outline potential research directions. This survey aims to foster understanding of this field, inspire new ideas, and encourage widespread applications in both academia and industry.
翻译:信息抽取(IE)旨在从非结构化文本中提取结构化信息,但常因数据稀缺和未见类别而在低资源场景中面临挑战。本文从传统方法和基于大语言模型(LLM)两个视角,综述了面向低资源信息抽取的神经方法,并将其系统归类为细粒度分类体系。随后,我们开展了基于LLM的方法与先前最优模型的实证研究,并发现:(1)经过良好微调的语言模型仍占主导地位;(2)对开源LLM进行微调及结合GPT系列进行上下文学习(ICL)总体上具有潜力;(3)低资源信息抽取的最优LLM技术方案可能因任务而异。此外,我们探讨了借助LLM的低资源信息抽取,着重介绍了有前景的应用,并概述了潜在研究方向。本综述旨在促进对该领域的理解、激发新思路,并推动其在学术界和工业界的广泛应用。