Electronic prior authorization workflows require FHIR Questionnaire items to carry LOINC codes, yet most items in the HL7 Da Vinci CDS-Library lack these bindings. We treat this as a retrieval problem: given a Questionnaire item's text, find the correct LOINC code in a pool of 97,314 active codes. We compare six methods (TF-IDF, frozen MiniLM, BioBERT, BioLORD, contrastively fine-tuned MiniLM, and a TF-IDF+GPT reranker) on a 54-item evaluation set spanning three query styles (natural question, medium, and terse). No single method wins on every metric. BioLORD, a frozen encoder pre-trained on biomedical ontology definitions, has the best top-rank accuracy (R@1 = 0.185, MRR = 0.246) despite seeing no task-specific data, while a contrastive fine-tune on raw LHC-Forms pairs takes R@5 (0.389) and R@10 (0.426). A distribution-shift ablation shows why the fine-tune in our main table is not the strongest one: adding GPT-generated paraphrases to the raw pairs drops R@5 from 0.389 to 0.296, so the augmented union underperforms raw-only training on every metric except R@1. Performance peaks at 5k training pairs. Error analysis on BioLORD's R@1 failures shows that wrong-specificity and ambiguous-text cases together account for 59% of errors.


翻译:电子预授权工作流要求FHIR问卷条目携带LOINC编码,但HL7 Da Vinci CDS-Library中的大部分条目缺乏此类绑定。我们将此问题视为检索任务:根据问卷条目的文本,从包含97,314个活跃编码的池中找出正确的LOINC编码。我们比较了六种方法(TF-IDF、冻结MiniLM、BioBERT、BioLORD、对比微调MiniLM及TF-IDF+GPT重排序器)在涵盖三种查询风格(自然问题、中等长度和简洁形式)的54项评估集上的表现。没有单一方法在所有指标上胜出。BioLORD作为基于生物医学本体定义预训练的冻结编码器,尽管未接触任务特定数据,仍取得了最佳顶级准确率(R@1=0.185,MRR=0.246),而在原始LHC-Forms对上进行对比微调的方法则在R@5(0.389)和R@10(0.426)上领先。分布偏移消融实验揭示了主表中微调模型并非最优的原因:在原始对中加入GPT生成的释义后,R@5从0.389降至0.296,因此除R@1外,增强联合训练在各项指标上均低于仅使用原始数据的训练方法。模型性能在5,000个训练对时达到峰值。对BioLORD在R@1失败案例的错误分析表明,特异性错误和歧义文本合计占错误总量的59%。

0
下载
关闭预览

相关内容

BERT/Transformer/迁移学习NLP资源大列表
专知
19+阅读 · 2019年6月9日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
迁移学习之Domain Adaptation
全球人工智能
18+阅读 · 2018年4月11日
行人再识别中的迁移学习
计算机视觉战队
11+阅读 · 2017年12月20日
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
13+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
3+阅读 · 今天14:49
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
相关VIP内容
相关基金
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
13+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
31+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员