Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to sequential, irreversible decisions under uncertainty. Static benchmarks cannot probe and existing interactive medical benchmarks each compromise on at least one of them. We present ClinEnv, an interactive benchmark that evaluates LLMs as attending physicians over real inpatient admissions under a paradigm we term Longitudinal Inpatient Simulation. Each case is automatically constructed into an ordered sequence of decision stages; at every stage the model must actively query four specialized agents before committing to medications, procedures, and diagnoses. ClinEnv scores both what the model decides, through deterministic ontology-grounded matching, and how it gathers information. Across seven models, the strongest reaches only 0.31 decision F1, and outcome quality is sharply decoupled from process quality. Difficulty concentrates in management decisions and later stages, where models recover discharge diagnoses far more reliably than management actions (0.51 vs. 0.17 F1) and continue to issue redundant queries as cases progress. ClinEnv makes this information-acquisition gap, invisible to outcome-only evaluation, directly measurable.


翻译:[翻译摘要] 临床实践并非从枚举选项中挑选答案:医生需逐步收集异质性信息,并在不确定性下做出顺序的、不可逆的决策。静态基准测试无法深入探究这一过程,而现有的交互式医学基准测试至少在其中某一方面存在妥协。我们提出ClinEnv——一个交互式基准测试,用于在一种我们称为"纵向住院模拟"的范式下,评估大语言模型作为主治医生处理真实住院病例的能力。每个病例被自动构建成有序的决策阶段序列;在每个阶段,模型必须在做出药物、程序和诊断决策之前,主动查询四个专业智能体。ClinEnv通过确定性的本体论基础匹配,对模型的决策结果以及信息收集过程进行评分。在七个模型的测试中,表现最佳的模型仅达到0.31的决策F1分数,且结果质量与过程质量严重脱钩。难度集中于管理决策及后期阶段——模型恢复出院诊断的可靠性远高于管理行动(F1值0.51 vs 0.17),且随着病例进展持续发出冗余查询。ClinEnv使这种仅关注结果的评估无法察觉的信息获取差距变得可直接量化。

0
下载
关闭预览

相关内容

Agentic RL:框架、实践与长程智能体训练
专知会员服务
23+阅读 · 6月24日
利用表示学习推动多机构电子健康记录数据研究
专知会员服务
16+阅读 · 2025年2月17日
Nature子刊综述|疾病建模和药物控制中的计算系统生物学
【AI与医学】多模态机器学习精准医疗健康
专知会员服务
83+阅读 · 2022年4月25日
【AI与医学】多模态机器学习精准医疗健康
医疗中的自动机器学习和可解释性
专知
24+阅读 · 2019年4月1日
NLP-Progress记录NLP最新数据集、论文和代码: 助你紧跟NLP前沿
中国人工智能学会
12+阅读 · 2018年11月15日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员