Electronic health records include information on patients' status and medical history, which could cover the history of diseases and disorders that could be hereditary. One important use of family history information is in precision health, where the goal is to keep the population healthy with preventative measures. Natural Language Processing (NLP) and machine learning techniques can assist with identifying information that could assist health professionals in identifying health risks before a condition is developed in their later years, saving lives and reducing healthcare costs. We survey the literature on the techniques from the NLP field that have been developed to utilise digital health records to identify risks of familial diseases. We highlight that rule-based methods are heavily investigated and are still actively used for family history extraction. Still, more recent efforts have been put into building neural models based on large-scale pre-trained language models. In addition to the areas where NLP has successfully been utilised, we also identify the areas where more research is needed to unlock the value of patients' records regarding data collection, task formulation and downstream applications.
翻译:电子健康记录包含患者状况和病史信息,其中可能涵盖具有遗传性的疾病史。家族史信息在精准健康领域具有重要应用价值——通过预防性措施维持人群健康。自然语言处理(NLP)与机器学习技术可协助识别关键信息,帮助医疗专业人员在其晚年疾病形成前识别健康风险,从而挽救生命并降低医疗成本。本文综述了NLP领域基于数字健康记录开发的技术文献,旨在识别家族性疾病风险。研究表明,基于规则的方法在家族史抽取领域得到了广泛研究且仍被积极采用,但近年来更多研究致力于构建基于大规模预训练语言模型的神经模型。除NLP已成功应用的领域外,本文还指出了在数据收集、任务制定和下游应用方面需要进一步研究的领域,以充分挖掘患者记录的价值。