Electronic health records (EHRs) hold significant value for research and applications. As a new way of information extraction, question answering (QA) can extract more flexible information than conventional methods and is more accessible to clinical researchers, but its progress is impeded by the scarcity of annotated data. In this paper, we propose a novel approach that automatically generates training data for transfer learning of QA models. Our pipeline incorporates a preprocessing module to handle challenges posed by extraction types that are not readily compatible with extractive QA frameworks, including cases with discontinuous answers and many-to-one relationships. The obtained QA model exhibits excellent performance on subtasks of information extraction in EHRs, and it can effectively handle few-shot or zero-shot settings involving yes-no questions. Case studies and ablation studies demonstrate the necessity of each component in our design, and the resulting model is deemed suitable for practical use.
翻译:电子健康记录(EHR)对研究与应用具有重要价值。作为一种新型信息提取方式,问答(QA)比传统方法能够提取更灵活的信息,且更便于临床研究人员使用,但其发展受限于标注数据的稀缺性。本文提出一种新型方法,可自动生成问答模型迁移学习的训练数据。我们的管道集成了预处理模块,用于处理与抽取式问答框架不兼容的提取类型所带来的挑战,包括答案不连续及多对一关系等情况。所获得的问答模型在电子病历信息提取子任务中表现出色,并能有效应对涉及是非问句的小样本或零样本场景。案例研究与消融实验证明了各设计组件的必要性,最终模型适用于实际应用场景。