Learning template based information extraction from documents is a crucial yet difficult task. Prior template-based IE approaches assume foreknowledge of the domain templates; however, real-world IE do not have pre-defined schemas and it is a figure-out-as you go phenomena. To quickly bootstrap templates in a real-world setting, we need to induce template slots from documents with zero or minimal supervision. Since the purpose of question answering intersect with the goal of information extraction, we use automatic question generation to induce template slots from the documents and investigate how a tiny amount of a proxy human-supervision on-the-fly (termed as InteractiveIE) can further boost the performance. Extensive experiments on biomedical and legal documents, where obtaining training data is expensive, reveal encouraging trends of performance improvement using InteractiveIE over AI-only baseline.
翻译:从文档中学习基于模板的信息抽取是一项关键但困难的任务。以往基于模板的信息抽取方法假定预知领域模板;然而,现实世界的信息抽取并无预定义模式,而是一个边发现边推进的过程。为在真实场景中快速启动模板,我们需要在零监督或极少监督条件下从文档中归纳模板槽位。鉴于问答目标与信息抽取目标存在交集,我们利用自动问题生成从文档中归纳模板槽位,并探究少量即时人工代理监督(称为InteractiveIE)如何进一步提升性能。在获取训练数据成本高昂的生物医学和法律文档上的大量实验表明,相较于仅依赖人工智能的基线方法,使用InteractiveIE展现出令人鼓舞的性能提升趋势。