Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to sequential, irreversible decisions under uncertainty. Static benchmarks cannot probe and existing interactive medical benchmarks each compromise on at least one of them. We present ClinEnv, an interactive benchmark that evaluates LLMs as attending physicians over real inpatient admissions under a paradigm we term Longitudinal Inpatient Simulation. Each case is automatically constructed into an ordered sequence of decision stages; at every stage the model must actively query four specialized agents before committing to medications, procedures, and diagnoses. ClinEnv scores both what the model decides, through deterministic ontology-grounded matching, and how it gathers information. Across seven models, the strongest reaches only 0.31 decision F1, and outcome quality is sharply decoupled from process quality. Difficulty concentrates in management decisions and later stages, where models recover discharge diagnoses far more reliably than management actions (0.51 vs. 0.17 F1) and continue to issue redundant queries as cases progress. ClinEnv makes this information-acquisition gap, invisible to outcome-only evaluation, directly measurable.
翻译:[翻译摘要]
临床实践并非从枚举选项中挑选答案:医生需逐步收集异质性信息,并在不确定性下做出顺序的、不可逆的决策。静态基准测试无法深入探究这一过程,而现有的交互式医学基准测试至少在其中某一方面存在妥协。我们提出ClinEnv——一个交互式基准测试,用于在一种我们称为"纵向住院模拟"的范式下,评估大语言模型作为主治医生处理真实住院病例的能力。每个病例被自动构建成有序的决策阶段序列;在每个阶段,模型必须在做出药物、程序和诊断决策之前,主动查询四个专业智能体。ClinEnv通过确定性的本体论基础匹配,对模型的决策结果以及信息收集过程进行评分。在七个模型的测试中,表现最佳的模型仅达到0.31的决策F1分数,且结果质量与过程质量严重脱钩。难度集中于管理决策及后期阶段——模型恢复出院诊断的可靠性远高于管理行动(F1值0.51 vs 0.17),且随着病例进展持续发出冗余查询。ClinEnv使这种仅关注结果的评估无法察觉的信息获取差距变得可直接量化。