Effectively representing medical concepts and patients is important for healthcare analytical applications. Representing medical concepts for healthcare analytical tasks requires incorporating medical domain knowledge and prior information from patient description data. Current methods, such as feature engineering and mapping medical concepts to standardized terminologies, have limitations in capturing the dynamic patterns from patient description data. Other embedding-based methods have difficulties in incorporating important medical domain knowledge and often require a large amount of training data, which may not be feasible for most healthcare systems. Our proposed framework, MD-Manifold, introduces a novel approach to medical concept and patient representation. It includes a new data augmentation approach, concept distance metric, and patient-patient network to incorporate crucial medical domain knowledge and prior data information. It then adapts manifold learning methods to generate medical concept-level representations that accurately reflect medical knowledge and patient-level representations that clearly identify heterogeneous patient cohorts. MD-Manifold also outperforms other state-of-the-art techniques in various downstream healthcare analytical tasks. Our work has significant implications in information systems research in representation learning, knowledge-driven machine learning, and using design science as middle-ground frameworks for downstream explorative and predictive analyses. Practically, MD-Manifold has the potential to create effective and generalizable representations of medical concepts and patients by incorporating medical domain knowledge and prior data information. It enables deeper insights into medical data and facilitates the development of new analytical applications for better healthcare outcomes.


翻译:有效表征医学概念和患者对于医疗分析应用至关重要。为医疗分析任务表征医学概念需要整合医学领域知识以及患者描述数据中的先验信息。当前方法(如特征工程和将医学概念映射至标准化术语)在捕捉患者描述数据的动态模式方面存在局限性。其他基于嵌入的方法难以纳入重要的医学领域知识,且通常需要大量训练数据,这对大多数医疗系统而言难以实现。我们提出的框架MD-Manifold引入了一种新颖的医学概念与患者表征方法,包含新的数据增强方法、概念距离度量以及患者-患者网络,以整合关键的医学领域知识和先验数据信息。随后,该方法适应流形学习技术生成准确反映医学知识的医学概念级表征,以及清晰识别异质性患者队列的患者级表征。MD-Manifold在各种下游医疗分析任务中的表现亦优于其他前沿技术。本研究对表征学习、知识驱动机器学习,以及将设计科学作为下游探索性与预测性分析中间框架的信息系统研究具有重大意义。在实际应用中,MD-Manifold通过整合医学领域知识与先验数据信息,有望创建有效且泛化的医学概念与患者表征,从而深入挖掘医疗数据洞察,推动开发新的分析应用以改善医疗成效。

0
下载
关闭预览

相关内容

LLM in Medical Domain: 大语言模型在医学领域的应用
专知会员服务
103+阅读 · 2023年6月17日
Nature Medicine | 多模态的生物医学AI
专知会员服务
31+阅读 · 2022年9月25日
Nat. Biomed. Eng.| 综述:医学和医疗保健中的自监督学习
专知会员服务
40+阅读 · 2022年8月25日
专知会员服务
68+阅读 · 2021年6月3日
因果关联学习,Causal Relational Learning
专知会员服务
185+阅读 · 2020年4月21日
MIT新书《强化学习与最优控制》
专知会员服务
283+阅读 · 2019年10月9日
多模态认知计算
专知
7+阅读 · 2022年9月16日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
无监督元学习表示学习
CreateAMind
27+阅读 · 2019年1月4日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
自然语言处理顶会EMNLP2018接受论文列表!
专知
87+阅读 · 2018年8月26日
【论文】图上的表示学习综述
机器学习研究会
15+阅读 · 2017年9月24日
国家自然科学基金
0+阅读 · 2016年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
Arxiv
69+阅读 · 2022年9月7日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
7+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
5+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
相关基金
国家自然科学基金
0+阅读 · 2016年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2013年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员