Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often privacy-sensitive. We developed a locally deployable framework that enables small language models to self-generate, verify, refine, and evaluate entity-specific prompts for extracting multiple clinical entities from dental notes. Using 1,200 annotated notes, we evaluated candidate open-weight models with multi-prompt ensemble inference and further adapted selected models using QLoRA-based supervised fine-tuning and direct preference optimization. Model performance varied substantially, highlighting the need for task-specific evaluation rather than reliance on generic benchmarks. Qwen2.5-14B-Instruct achieved the strongest baseline performance. After DPO, Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct achieved micro/macro F1 scores of 0.864/0.837 and 0.806/0.797, respectively. These findings suggest that automated prompt optimization combined with lightweight preference-based post-training can support scalable clinical information extraction using locally deployed small language models.


翻译:牙科病程记录的临床命名实体识别具有高度非结构化、领域特异性强及隐私敏感性高的特点,这使得该任务面临诸多挑战。我们开发了一种可本地部署的框架,使小语言模型能够自主生成、验证、优化及评估面向实体的提示,从而从牙科记录中抽取多种临床实体。基于1200条已标注记录,我们采用多提示集成推理对候选开源模型进行评估,并进一步通过基于QLoRA的监督微调与直接偏好优化对所选模型进行适配。模型性能差异显著,这表明需开展任务特定评估而非依赖通用基准。Qwen2.5-14B-Instruct取得了最优基线性能。经DPO优化后,Qwen2.5-14B-Instruct与Llama-3.1-8B-Instruct的宏/微平均F1值分别达到0.864/0.837及0.806/0.797。研究结果表明,自动提示优化结合轻量级偏好后训练能够支撑基于本地部署小语言模型的可扩展临床信息抽取。

0
下载
关闭预览

相关内容

知识抽取,即从不同来源、不同结构的数据中进行知识提取,形成知识(结构化数据)存入到知识图谱。
医学领域大型语言模型的新进展
专知会员服务
25+阅读 · 2025年10月5日
小型语言模型综述
专知会员服务
56+阅读 · 2024年10月29日
大语言模型中的提示隐私保护
专知会员服务
24+阅读 · 2024年7月24日
LLM in Medical Domain: 大语言模型在医学领域的应用
专知会员服务
103+阅读 · 2023年6月17日
专知会员服务
40+阅读 · 2021年5月14日
专知会员服务
126+阅读 · 2021年4月29日
专知会员服务
52+阅读 · 2021年3月28日
小样本学习(Few-shot Learning)综述
云栖社区
22+阅读 · 2019年4月6日
一种关键字提取新方法
1号机器人网
21+阅读 · 2018年11月15日
TextInfoExp:自然语言处理相关实验(基于sougou数据集)
全球人工智能
12+阅读 · 2017年11月12日
NLP中自动生产文摘(auto text summarization)
机器学习研究会
14+阅读 · 2017年10月10日
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
1+阅读 · 今天14:36
《战争中的大语言模型监管》
专知会员服务
2+阅读 · 今天14:26
边缘计算的军事应用
专知会员服务
8+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
相关VIP内容
医学领域大型语言模型的新进展
专知会员服务
25+阅读 · 2025年10月5日
小型语言模型综述
专知会员服务
56+阅读 · 2024年10月29日
大语言模型中的提示隐私保护
专知会员服务
24+阅读 · 2024年7月24日
LLM in Medical Domain: 大语言模型在医学领域的应用
专知会员服务
103+阅读 · 2023年6月17日
专知会员服务
40+阅读 · 2021年5月14日
专知会员服务
126+阅读 · 2021年4月29日
专知会员服务
52+阅读 · 2021年3月28日
相关基金
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员