Information extraction(IE) is a crucial subfield within natural language processing. In this study, we introduce a Sentence Classification and Named Entity Recognition Multi-task (SCNM) approach that combines Sentence Classification (SC) and Named Entity Recognition (NER). We develop a Sentence-to-Label Generation (SLG) framework for SCNM and construct a Wikipedia dataset containing both SC and NER. Using a format converter, we unify input formats and employ a generative model to generate SC-labels, NER-labels, and associated text segments. We propose a Constraint Mechanism (CM) to improve generated format accuracy. Our results show SC accuracy increased by 1.13 points and NER by 1.06 points in SCNM compared to standalone tasks, with CM raising format accuracy from 63.61 to 100. The findings indicate mutual reinforcement effects between SC and NER, and integration enhances both tasks' performance.
翻译:信息抽取(IE)是自然语言处理中的重要子领域。本研究提出一种结合句子分类(SC)与命名实体识别(NER)的多任务方法——SCNM(句子分类与命名实体识别多任务方法)。我们为SCNM开发了句子到标签生成(SLG)框架,并构建了同时包含SC与NER标注的维基百科数据集。通过格式转换器统一输入格式,采用生成式模型同步生成SC标签、NER标签及其对应的文本片段。我们提出约束机制(CM)提升生成格式的准确性。实验结果表明:相较于独立任务,SCNM方法使SC准确率提升1.13个百分点,NER准确率提升1.06个百分点;CM将格式准确率从63.61%提升至100%。研究证实SC与NER存在相互增强效应,其联合建模能够显著提升两个任务的性能表现。