Textual health records of cancer patients are usually protracted and highly unstructured, making it very time-consuming for health professionals to get a complete overview of the patient's therapeutic course. As such limitations can lead to suboptimal and/or inefficient treatment procedures, healthcare providers would greatly benefit from a system that effectively summarizes the information of those records. With the advent of deep neural models, this objective has been partially attained for English clinical texts, however, the research community still lacks an effective solution for languages with limited resources. In this paper, we present the approach we developed to extract procedures, drugs, and diseases from oncology health records written in European Portuguese. This project was conducted in collaboration with the Portuguese Institute for Oncology which, besides holding over $10$ years of duly protected medical records, also provided oncologist expertise throughout the development of the project. Since there is no annotated corpus for biomedical entity extraction in Portuguese, we also present the strategy we followed in annotating the corpus for the development of the models. The final models, which combined a neural architecture with entity linking, achieved $F_1$ scores of $88.6$, $95.0$, and $55.8$ per cent in the mention extraction of procedures, drugs, and diseases, respectively.
翻译:癌症患者的文本健康记录通常冗长且高度非结构化,这使得医疗专业人员难以快速全面掌握患者的治疗过程。由于此类限制可能导致次优和/或低效的治疗方案,医疗机构将极大受益于能够有效总结这些记录信息的系统。随着深度神经模型的出现,这一目标在英文临床文本中已部分实现,然而,研究界仍缺乏针对资源有限语言的有效解决方案。本文介绍了我们开发的从欧洲葡萄牙语肿瘤健康记录中抽取手术、药物和疾病的方法。该项目与葡萄牙肿瘤研究所合作开展,该机构不仅拥有超过10年的受保护医疗记录,还在项目开发过程中提供了肿瘤学专家的专业知识。由于目前尚无用于葡萄牙语生物医学实体抽取的标注语料库,我们还介绍了为开发模型而标注语料库所采用的策略。最终模型将神经架构与实体链接相结合,在手术、药物和疾病的提及抽取中分别达到了$88.6$%、$95.0$%和$55.8$%的$F_1$分数。