Large language models like GPT-3.5-turbo and GPT-4 hold promise for healthcare professionals, but they may inadvertently inherit biases during their training, potentially affecting their utility in medical applications. Despite few attempts in the past, the precise impact and extent of these biases remain uncertain. Through both qualitative and quantitative analyses, we find that these models tend to project higher costs and longer hospitalizations for White populations and exhibit optimistic views in challenging medical scenarios with much higher survival rates. These biases, which mirror real-world healthcare disparities, are evident in the generation of patient backgrounds, the association of specific diseases with certain races, and disparities in treatment recommendations, etc. Our findings underscore the critical need for future research to address and mitigate biases in language models, especially in critical healthcare applications, to ensure fair and accurate outcomes for all patients.
翻译:像GPT-3.5-turbo和GPT-4这样的大型语言模型为医疗专业人员带来了希望,但它们可能在训练过程中无意中继承了偏见,这可能影响其在医学应用中的效用。尽管过去有少数尝试,但这些偏见的确切影响和程度仍不确定。通过定性和定量分析,我们发现这些模型倾向于为白人群体预估更高的医疗费用和更长的住院时间,并在具有挑战性的医学场景中表现出乐观态度,给出更高的存活率。这些偏见反映了现实世界中的医疗保健差异,体现在患者背景生成、特定疾病与特定种族的关联以及治疗建议的差异等方面。我们的发现强调了未来研究在解决和减轻语言模型偏见方面的关键需求,尤其是在关键的医疗保健应用中,以确保所有患者都能获得公平和准确的结果。