Pain is a common reason for accessing healthcare resources and is a growing area of research, especially in its overlap with mental health. Mental health electronic health records are a good data source to study this overlap. However, much information on pain is held in the free text of these records, where mentions of pain present a unique natural language processing problem due to its ambiguous nature. This project uses data from an anonymised mental health electronic health records database. The data are used to train a machine learning based classification algorithm to classify sentences as discussing patient pain or not. This will facilitate the extraction of relevant pain information from large databases, and the use of such outputs for further studies on pain and mental health. 1,985 documents were manually triple-annotated for creation of gold standard training data, which was used to train three commonly used classification algorithms. The best performing model achieved an F1-score of 0.98 (95% CI 0.98-0.99).
翻译:疼痛是获取医疗资源的常见原因,也是研究不断增长的领域,尤其是在其与心理健康的重叠方面。心理健康电子健康记录是研究这种重叠的良好数据来源。然而,大量疼痛信息存储在这些记录的自由文本中,其中疼痛的提及因其模糊性构成了独特的自然语言处理问题。本项目使用来自匿名化心理健康电子健康记录数据库的数据。这些数据用于训练基于机器学习的分类算法,以判断句子是否讨论患者的疼痛。这将有助于从大型数据库中提取相关疼痛信息,并利用此类输出进行关于疼痛与心理健康的进一步研究。共对1985份文档进行了人工三重标注,以创建黄金标准训练数据,并用于训练三种常用分类算法。表现最佳的模型F1分数达到0.98(95%置信区间0.98-0.99)。