Traditional Korean medicine (TKM) emphasizes individualized diagnosis and treatment. This uniqueness makes AI modeling difficult due to limited data and implicit processes. Large language models (LLMs) have demonstrated impressive medical inference, even without advanced training in medical texts. This study assessed the capabilities of GPT-4 in TKM, using the Korean National Licensing Examination for Korean Medicine Doctors (K-NLEKMD) as a benchmark. The K-NLEKMD, administered by a national organization, encompasses 12 major subjects in TKM. We optimized prompts with Chinese-term annotation, English translation for questions and instruction, exam-optimized instruction, and self-consistency. GPT-4 with optimized prompts achieved 66.18% accuracy, surpassing both the examination's average pass mark of 60% and the 40% minimum for each subject. The gradual introduction of language-related prompts and prompting techniques enhanced the accuracy from 51.82% to its maximum accuracy. GPT-4 showed low accuracy in subjects including public health & medicine-related law, internal medicine (2) which are localized in Korea and TKM. The model's accuracy was lower for questions requiring TKM-specialized knowledge. It exhibited higher accuracy in diagnosis-based and recall-based questions than in intervention-based questions. A positive correlation was observed between the consistency and accuracy of GPT-4's responses. This study unveils both the potential and challenges of applying LLMs to TKM. These findings underline the potential of LLMs like GPT-4 in culturally adapted medicine, especially TKM, for tasks such as clinical assistance, medical education, and research. But they also point towards the necessity for the development of methods to mitigate cultural bias inherent in large language models and validate their efficacy in real-world clinical settings.
翻译:传统韩医学(TKM)强调个体化诊断与治疗,其独特性因数据有限及隐式过程而给人工智能建模带来挑战。大型语言模型(LLM)即使在未经过医学文本高级训练的情况下,也展现出令人瞩目的医学推理能力。本研究以韩国韩医医师国家执业考试(K-NLEKMD)为基准,评估了GPT-4在TKM领域的能力。K-NLEKMD由国家级机构管理,涵盖TKM的12个主要学科。我们通过中文术语注释、题目及指令的英文翻译、考试优化指令以及自一致性方法优化了提示词。采用优化提示词的GPT-4准确率达到66.18%,分别超过考试平均及格线(60%)及各科最低分(40%)的要求。语言相关提示词与提示技术的渐进引入将准确率从51.82%提升至最高水平。GPT-4在公共卫生与医学相关法律、内科学(2)等韩国本土化及TKM特有科目中准确率较低;模型对需要TKM专业知识的题目表现较差,在基于诊断与基于回忆的题目中准确率高于基于干预的题目。GPT-4回答的一致性与准确率呈正相关。本研究揭示了将LLM应用于TKM的潜力与挑战,结果表明GPT-4等LLM在文化适应性医学(尤其TKM)的临床辅助、医学教育与研究等任务中具有潜力,但也指出了需开发方法以减轻大型语言模型固有文化偏见、并验证其在真实临床场景中效能的必要性。