The use of Electronic Health Records (EHRs) has increased dramatically in the past 15 years, as, it is considered an important source of managing data od patients. The EHRs are primary sources of disease diagnosis and demographic data of patients worldwide. Therefore, the data can be utilized for secondary tasks such as research. This paper aims to make such data usable for research activities such as monitoring disease statistics for a specific population. As a result, the researchers can detect the disease causes for the behavior and lifestyle of the target group. One of the limitations of EHRs systems is that the data is not available in the standard format but in various forms. Therefore, it is required to first convert the names of the diseases and demographics data into one standardized form to make it usable for research activities. There is a large amount of EHRs available, and solving the standardizing issues requires some optimized techniques. We used a first-hand EHR dataset extracted from EHR systems. Our application uploads the dataset from the EHRs and converts it to the ICD-10 coding system to solve the standardization problem. So, we first apply the steps of pre-processing, annotation, and transforming the data to convert it into the standard form. The data pre-processing is applied to normalize demographic formats. In the annotation step, a machine learning model is used to recognize the diseases from the text. Furthermore, the transforming step converts the disease name to the ICD-10 coding format. The model was evaluated manually by comparing its performance in terms of disease recognition with an available dictionary-based system (MetaMap). The accuracy of the proposed machine learning model is 81%, that outperformed MetaMap accuracy of 67%. This paper contributed to system modelling for EHR data extraction to support research activities.
翻译:过去15年间,电子健康记录(EHRs)的使用急剧增加,因其被视为患者数据管理的重要来源。EHRs是全球范围内疾病诊断和患者人口统计数据的主要来源,因此这些数据可用于研究等辅助任务。本文旨在使此类数据能够用于监测特定人群疾病统计等研究活动,使研究人员能够检测目标群体行为与生活方式所导致的疾病原因。EHR系统的一个局限性在于数据并非以标准格式呈现,而是以多种形式存在。因此,首先需要将疾病名称和人口统计数据转换为统一标准形式,使其可用于研究活动。现有大量EHR数据,解决标准化问题需要优化技术。我们使用从EHR系统直接提取的一手EHR数据集。应用程序从EHRs上传数据集,并将其转换为ICD-10编码系统以解决标准化问题。为此,我们首先执行数据预处理、标注和转换步骤,将数据转化为标准形式。数据预处理用于规范人口统计格式;标注步骤中,使用机器学习模型从文本中识别疾病;转换步骤则将疾病名称转换为ICD-10编码格式。通过将模型在疾病识别方面的性能与基于字典的系统(MetaMap)进行手动对比评估,本机器学习模型的准确率为81%,优于MetaMap的67%。本文为支持研究活动的EHR数据提取的系统建模做出了贡献。