Emergency department triage assigns patients an acuity score that determines treatment priority, and clinical evidence documents persistent gender disparities in human acuity assessment. As hospitals pilot large language models (LLMs) as triage decision support, a critical question is whether these models reproduce or mitigate known biases. We present EQUITRIAGE, a fairness audit of LLM-based ESI assignment evaluating five models (Gemini-3-Flash, Nemotron-3-Super, DeepSeek-V3.1, Mistral-Small-3.2, GPT-4.1-Nano) across 374,275 evaluations on 18,714 MIMIC-IV-ED vignettes under four prompt strategies. Of 9,368 originals, 9,346 are paired with a gender-swapped counterfactual. All five models produced flip rates above a pre-registered 5% threshold (9.9% to 43.8%). Two showed directional female undertriage (DeepSeek F/M 2.15:1, Gemini 1.34:1); two were near-parity; one had high sensitivity with weak male-direction asymmetry. DeepSeek's directional bias coexisted with a low outcome-linked calibration gap (0.013 against MIMIC-IV admission), a Chouldechova-style dissociation between within-group calibration and between-pair counterfactual invariance. Demographic blinding reduced Gemini's flip rate to 0.5%; an age-preserving blind variant left DeepSeek with residual F/M 1.25, implicating age as a residual channel. Chain-of-thought prompting degraded accuracy for all five models. A two-model ablation reveals opposite underlying mechanisms for the same directional phenotype: in Gemini the signal is emergent in the combined name+gender swap, while in DeepSeek the gender token alone carries it. EQUITRIAGE shows that group parity, counterfactual invariance, and gender calibration are distinct fairness properties, that intervention effectiveness is model-dependent, and that per-model counterfactual auditing should precede clinical deployment.


翻译:急诊分诊通过为患者分配急症评分以确定治疗优先级,而临床证据表明,在人工急症评估中持续存在性别差异。随着医院试点将大语言模型(LLM)作为分诊决策支持工具,一个关键问题在于这些模型是再现还是减轻了已知偏见。我们提出EQUITRIAGE,一项对基于LLM的急诊严重度指数(ESI)分配的公平性审计,评估了五种模型(Gemini-3-Flash、Nemotron-3-Super、DeepSeek-V3.1、Mistral-Small-3.2、GPT-4.1-Nano),在四种提示策略下对18,714个MIMIC-IV-ED病例场景进行了374,275次评估。在9,368个原始场景中,有9,346个与性别交换的反事实场景配对。所有五种模型的翻转率均超过预注册的5%阈值(9.9%至43.8%)。两种模型表现出定向的女性分诊不足(DeepSeek女性/男性比例2.15:1,Gemini 1.34:1);两种模型接近无差异;一种模型具有高敏感性但存在微弱的男性方向不对称性。DeepSeek的定向偏见与较低的结果相关校准差距(相对于MIMIC-IV入院率0.013)并存,呈现出Chouldechova式的群体内校准与配对间反事实不变性之间的分离。人口统计信息屏蔽将Gemini的翻转率降至0.5%;保留年龄的屏蔽变体使DeepSeek仍存在残差女性/男性比例1.25,表明年龄是残差传递通道。思维链提示降低了所有五种模型的准确性。一项双模型消融实验揭示了相同定向表现型背后的相反机制:在Gemini中,信号出现在综合姓名+性别交换中,而在DeepSeek中,仅性别标记本身即可携带该信号。EQUITRIAGE表明,群体平等性、反事实不变性和性别校准是不同的公平性属性,干预效果具有模型依赖性,且临床部署前应针对每个模型进行反事实审计。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
智能体评判者(Agent-as-a-Judge)研究综述
专知会员服务
37+阅读 · 1月9日
大型语言模型中隐性与显性偏见的综合研究
专知会员服务
17+阅读 · 2025年11月25日
LLM/智能体作为数据分析师:综述
专知会员服务
38+阅读 · 2025年9月30日
迈向LLM时代的可泛化评估:超越基准的综述
专知会员服务
23+阅读 · 2025年4月29日
大型语言模型公平性
专知会员服务
41+阅读 · 2023年8月31日
异常检测(Anomaly Detection)综述
极市平台
20+阅读 · 2020年10月24日
情感计算综述
人工智能学家
34+阅读 · 2019年4月6日
异常检测的阈值,你怎么选?给你整理好了...
机器学习算法与Python学习
10+阅读 · 2018年9月19日
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
9+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
14+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
相关基金
国家自然科学基金
23+阅读 · 2016年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员