Emergency department triage assigns patients an acuity score that determines treatment priority, and clinical evidence documents persistent gender disparities in human acuity assessment. As hospitals pilot large language models (LLMs) as triage decision support, a critical question is whether these models reproduce or mitigate known biases. We present EQUITRIAGE, a fairness audit of LLM-based ESI assignment evaluating five models (Gemini-3-Flash, Nemotron-3-Super, DeepSeek-V3.1, Mistral-Small-3.2, GPT-4.1-Nano) across 374,275 evaluations on 18,714 MIMIC-IV-ED vignettes under four prompt strategies. Of 9,368 originals, 9,346 are paired with a gender-swapped counterfactual. All five models produced flip rates above a pre-registered 5% threshold (9.9% to 43.8%). Two showed directional female undertriage (DeepSeek F/M 2.15:1, Gemini 1.34:1); two were near-parity; one had high sensitivity with weak male-direction asymmetry. DeepSeek's directional bias coexisted with a low outcome-linked calibration gap (0.013 against MIMIC-IV admission), a Chouldechova-style dissociation between within-group calibration and between-pair counterfactual invariance. Demographic blinding reduced Gemini's flip rate to 0.5%; an age-preserving blind variant left DeepSeek with residual F/M 1.25, implicating age as a residual channel. Chain-of-thought prompting degraded accuracy for all five models. A two-model ablation reveals opposite underlying mechanisms for the same directional phenotype: in Gemini the signal is emergent in the combined name+gender swap, while in DeepSeek the gender token alone carries it. EQUITRIAGE shows that group parity, counterfactual invariance, and gender calibration are distinct fairness properties, that intervention effectiveness is model-dependent, and that per-model counterfactual auditing should precede clinical deployment.
翻译:急诊分诊通过为患者分配急症评分以确定治疗优先级,而临床证据表明,在人工急症评估中持续存在性别差异。随着医院试点将大语言模型(LLM)作为分诊决策支持工具,一个关键问题在于这些模型是再现还是减轻了已知偏见。我们提出EQUITRIAGE,一项对基于LLM的急诊严重度指数(ESI)分配的公平性审计,评估了五种模型(Gemini-3-Flash、Nemotron-3-Super、DeepSeek-V3.1、Mistral-Small-3.2、GPT-4.1-Nano),在四种提示策略下对18,714个MIMIC-IV-ED病例场景进行了374,275次评估。在9,368个原始场景中,有9,346个与性别交换的反事实场景配对。所有五种模型的翻转率均超过预注册的5%阈值(9.9%至43.8%)。两种模型表现出定向的女性分诊不足(DeepSeek女性/男性比例2.15:1,Gemini 1.34:1);两种模型接近无差异;一种模型具有高敏感性但存在微弱的男性方向不对称性。DeepSeek的定向偏见与较低的结果相关校准差距(相对于MIMIC-IV入院率0.013)并存,呈现出Chouldechova式的群体内校准与配对间反事实不变性之间的分离。人口统计信息屏蔽将Gemini的翻转率降至0.5%;保留年龄的屏蔽变体使DeepSeek仍存在残差女性/男性比例1.25,表明年龄是残差传递通道。思维链提示降低了所有五种模型的准确性。一项双模型消融实验揭示了相同定向表现型背后的相反机制:在Gemini中,信号出现在综合姓名+性别交换中,而在DeepSeek中,仅性别标记本身即可携带该信号。EQUITRIAGE表明,群体平等性、反事实不变性和性别校准是不同的公平性属性,干预效果具有模型依赖性,且临床部署前应针对每个模型进行反事实审计。