The nomenclature of human disease has developed organically over the past centuries using Greek, Latin, and Arabic terminology and reflects the idiosyncrasies of different eras of medical discovery. Despite evident heterogeneity in naming practices, no systematic framework exists for characterising these conventions across all diseases. In this paper, we describe the Nomenclature Ontology for Medical And Disease names (NOMAD), a meta-taxonomy that classifies disease names according to their naming conventions. We developed a two-level taxonomy comprising 9 top-level categories and 20 subcategories and applied it to 22,548 index entries from the ICD-10-CM 2026 Alphabetical Index in a scalable three-stage machine learning-driven classification pipeline. Classification was multi-label, reflecting the compositional nature of medical nomenclature. We classified 99.1% of terms with a mean of 2.12 labels per entry. Anatomical categories were the most prevalent (63.8% of entries), followed by Descriptive (48.4%) and Pathophysiological (40.2%), while Eponymous and Geographical labels were less common than their cultural prominence might suggest (9.7% and 1.9% respectively). Among all Eponymous diseases, we identified only 57 (2.6%) of diseases named after a female person. We manually reviewed a random sample of n=2,255 entries (10%) for accuracy and calculated a full agreement rate of 70% and partial agreement rate of 26% (macro-averaged Cohen's Kappa score 0.832). Naming convention profiles varied substantially across ICD-10-CM chapters, reflecting specialty-specific epistemological traditions: infectious disease chapters were dominated by etiological labels and showed the highest proportion of geographical region related labels, the circulatory chapter by anatomical and pathophysiological labels, and mental and behavioural disorders showed the highest prevalence of socio-behavioral labels.


翻译:人类疾病命名体系在过去几个世纪中,借助希腊语、拉丁语和阿拉伯语术语自然演化,折射出不同医学发现时代的特有印记。尽管命名实践存在显著异质性,目前仍缺乏系统化框架来统括所有疾病的命名惯例。本文提出NOMAD(医学与疾病名称本体论),一种按命名惯例对疾病名称进行元分类的体系。我们构建了包含9个顶层类别和20个子类别的双层分类法,并在可扩展的三阶段机器学习驱动分类流水线中,将其应用于ICD-10-CM 2026字母索引中22,548条索引条目。分类采用多标签模式以反映医学命名的复合性特征,共完成99.1%的术语分类,平均每项条目包含2.12个标签。其中解剖学类别最常见(占条目63.8%),其次为描述性(48.4%)和病理生理学(40.2%),而基于人名和地理标签的使用频率低于其文化显著性所暗示的程度(分别占9.7%和1.9%)。在所有以人名命名的疾病中,仅发现57种(占2.6%)以女性命名。我们对随机选取的n=2,255条条目(10%)进行人工准确性审核,得到完全一致率为70%,部分一致率为26%(宏观平均Cohen's Kappa系数0.832)。各ICD-10-CM章节的命名惯例特征存在显著差异,反映专科特有的认识论传统:传染病章节以病因学标签为主且地理区域相关标签比例最高,循环系统章节以解剖学和病理生理学标签为主,而精神与行为障碍则呈现最高的社会行为学标签占比。

0
下载
关闭预览

相关内容

「中文电子病历命名实体识别」的研究与进展
专知会员服务
31+阅读 · 2022年11月5日
基于多来源文本的中文医学知识图谱的构建
专知会员服务
53+阅读 · 2020年8月21日
【NER综述】近五年中文电子病历命名实体识别研究进展
深度学习自然语言处理
12+阅读 · 2020年8月24日
基于多来源文本的中文医学知识图谱的构建
一文读懂命名实体识别
人工智能头条
33+阅读 · 2019年3月29日
本体:一文读懂领域本体构建
AINLP
40+阅读 · 2019年2月27日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
11+阅读 · 7月19日
相关VIP内容
「中文电子病历命名实体识别」的研究与进展
专知会员服务
31+阅读 · 2022年11月5日
基于多来源文本的中文医学知识图谱的构建
专知会员服务
53+阅读 · 2020年8月21日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员