This paper presents a question-answering approach to extract document-level event-argument structures. We automatically ask and answer questions for each argument type an event may have. Questions are generated using manually defined templates and generative transformers. Template-based questions are generated using predefined role-specific wh-words and event triggers from the context document. Transformer-based questions are generated using large language models trained to formulate questions based on a passage and the expected answer. Additionally, we develop novel data augmentation strategies specialized in inter-sentential event-argument relations. We use a simple span-swapping technique, coreference resolution, and large language models to augment the training instances. Our approach enables transfer learning without any corpora-specific modifications and yields competitive results with the RAMS dataset. It outperforms previous work, and it is especially beneficial to extract arguments that appear in different sentences than the event trigger. We also present detailed quantitative and qualitative analyses shedding light on the most common errors made by our best model.
翻译:本文提出了一种基于问答的方法,用于抽取文档级事件-论元结构。我们针对事件可能具有的每种论元类型自动提问并回答。问题通过手动定义的模板和生成式Transformer生成。基于模板的问题利用预定义的与角色相关的疑问词和上下文文档中的事件触发词生成。基于Transformer的问题则使用经过训练的大语言模型,根据段落和预期答案来构建问题。此外,我们开发了专门针对跨句事件-论元关系的新型数据增强策略,采用简单的片段交换技术、共指消解和大语言模型来扩充训练实例。该方法无需针对特定语料库进行修改即可实现迁移学习,并在RAMS数据集上取得了具有竞争力的结果。其性能优于之前的工作,尤其有利于抽取与事件触发词出现在不同句子中的论元。我们还提供了详细的定量与定性分析,揭示了最佳模型最常出现的错误类型。