The potential of large language models in medicine for education and decision making purposes has been demonstrated as they achieve decent scores on medical exams such as the United States Medical Licensing Exam (USMLE) and the MedQA exam. In this work, we evaluate the performance of ChatGPT-4 in the specialized field of radiation oncology using the 38th American College of Radiology (ACR) radiation oncology in-training (TXIT) exam and the 2022 red journal gray zone cases. For the TXIT exam, ChatGPT-3.5 and ChatGPT-4 have achieved the scores of 63.65% and 74.57%, respectively, highlighting the advantage of the latest ChatGPT-4 model. Based on the TXIT exam, ChatGPT-4's strong and weak areas in radiation oncology are identified to some extent. Specifically, ChatGPT-4 demonstrates good knowledge of statistics, CNS & eye, pediatrics, biology, and physics but has limitations in bone & soft tissue and gynecology, as per the ACR knowledge domain. Regarding clinical care paths, ChatGPT-4 performs well in diagnosis, prognosis, and toxicity but lacks proficiency in topics related to brachytherapy and dosimetry, as well as in-depth questions from clinical trials. For the gray zone cases, ChatGPT-4 is able to suggest a personalized treatment approach to each case and achieves comparable votes (28.76% on average) to human experts in general. Most importantly, it provides complementary suggestion to the recommendation from a single expert. Both evaluations have demonstrated the potential of ChatGPT in medical education for the general public and cancer patients, as well as the potential to aid clinical decision-making, while acknowledging its limitations in certain domains.
翻译:大型语言模型在医学教育与决策方面的潜力已得到证实,其在美国医学执照考试(USMLE)和MedQA等医学考试中取得了可观成绩。本研究利用第38届美国放射学会(ACR)放射肿瘤学住院医师培训(TXIT)考试及2022年《红皮杂志》灰色地带病例,评估ChatGPT-4在放射肿瘤学专业领域的表现。在TXIT考试中,ChatGPT-3.5和ChatGPT-4分别获得63.65%和74.57%的分数,凸显了最新ChatGPT-4模型的技术优势。基于TXIT考试,本研究在一定程度上识别了ChatGPT-4在放射肿瘤学中的优势与薄弱领域。具体而言,根据ACR知识领域划分,ChatGPT-4在统计学、中枢神经系统与眼科、儿科、生物学及物理学方面表现出色,但在骨与软组织及妇科领域存在局限。在临床诊疗路径方面,ChatGPT-4在诊断、预后和毒性方面表现良好,但在近距离治疗与剂量学相关课题及来自临床试验的深度问题方面能力不足。针对灰色地带病例,ChatGPT-4能为每个病例提出个性化治疗方案,其投票结果(平均获得28.76%支持率)总体上与人类专家相当。更重要的是,它能对单一位专家的建议提供补充性意见。两项评估均表明ChatGPT在面向公众及癌症患者的医学教育领域具有潜力,并可能辅助临床决策,同时也承认其在某些领域的局限性。