Despite growing interest in using large language models (LLMs) in healthcare, current explorations do not assess the real-world utility and safety of LLMs in clinical settings. Our objective was to determine whether two LLMs can serve information needs submitted by physicians as questions to an informatics consultation service in a safe and concordant manner. Sixty six questions from an informatics consult service were submitted to GPT-3.5 and GPT-4 via simple prompts. 12 physicians assessed the LLM responses' possibility of patient harm and concordance with existing reports from an informatics consultation service. Physician assessments were summarized based on majority vote. For no questions did a majority of physicians deem either LLM response as harmful. For GPT-3.5, responses to 8 questions were concordant with the informatics consult report, 20 discordant, and 9 were unable to be assessed. There were 29 responses with no majority on "Agree", "Disagree", and "Unable to assess". For GPT-4, responses to 13 questions were concordant, 15 discordant, and 3 were unable to be assessed. There were 35 responses with no majority. Responses from both LLMs were largely devoid of overt harm, but less than 20% of the responses agreed with an answer from an informatics consultation service, responses contained hallucinated references, and physicians were divided on what constitutes harm. These results suggest that while general purpose LLMs are able to provide safe and credible responses, they often do not meet the specific information need of a given question. A definitive evaluation of the usefulness of LLMs in healthcare settings will likely require additional research on prompt engineering, calibration, and custom-tailoring of general purpose models.
翻译:尽管大型语言模型在医疗领域的应用日益受到关注,但当前探索尚未评估其在临床场景中的现实效用与安全性。本研究旨在确定两类大型语言模型能否以安全且一致的方式,满足医生通过信息学咨询服务平台提交的信息需求。我们将66个源自信息学咨询服务的问题,通过简单提示提交至GPT-3.5与GPT-4模型。12名医生评估了模型回复可能对患者造成的伤害程度,以及其与现有信息学咨询报告的一致性。医生评估结果基于多数投票原则汇总。针对所有问题,多数医生均未判定任一模型的回复具有危害性。对于GPT-3.5,其回复中与信息学咨询报告一致的共8项、不一致的20项、无法评估的9项;另有29项回复在“同意”“不同意”“无法评估”三项上未形成多数意见。对于GPT-4,一致回复13项、不一致15项、无法评估3项;35项回复未形成多数意见。两类模型的回复基本未出现明显危害,但仅有不到20%的回复与信息学咨询服务的答案一致;回复中存在虚构引用,且医生对危害的定义存在分歧。结果表明,虽然通用型大型语言模型能提供安全且可信的回复,但往往无法满足特定问题的具体信息需求。要全面评估大型语言模型在医疗场景中的实用性,可能需进一步研究提示工程、校准方法及通用模型的定制化调优。