Recent investigations show that large language models (LLMs), specifically GPT-4, not only have remarkable capabilities in common Natural Language Processing (NLP) tasks but also exhibit human-level performance on various professional and academic benchmarks. However, whether GPT-4 can be directly used in practical applications and replace traditional artificial intelligence (AI) tools in specialized domains requires further experimental validation. In this paper, we explore the potential of LLMs such as GPT-4 to outperform traditional AI tools in dementia diagnosis. Comprehensive comparisons between GPT-4 and traditional AI tools are conducted to examine their diagnostic accuracy in a clinical setting. Experimental results on two real clinical datasets show that, although LLMs like GPT-4 demonstrate potential for future advancements in dementia diagnosis, they currently do not surpass the performance of traditional AI tools. The interpretability and faithfulness of GPT-4 are also evaluated by comparison with real doctors. We discuss the limitations of GPT-4 in its current state and propose future research directions to enhance GPT-4 in dementia diagnosis.
翻译:近期研究表明,大语言模型(LLMs),特别是GPT-4,不仅在常见自然语言处理(NLP)任务中展现出卓越能力,还在各类专业和学术基准测试中达到人类级水平。然而,GPT-4能否直接应用于实际场景并取代专业领域的传统人工智能(AI)工具,仍需进一步的实验验证。本文探索了GPT-4等大语言模型在痴呆症诊断中超越传统AI工具的潜力。通过全面比较GPT-4与传统AI工具在临床环境中的诊断准确性,基于两个真实临床数据集的实验结果显示:尽管GPT-4等大语言模型在痴呆症诊断领域展现出未来发展的潜力,但当前阶段其性能尚未超越传统AI工具。通过与真实医生的对比,我们还评估了GPT-4的可解释性与忠实度。本文讨论了GPT-4当前存在的局限性,并提出了未来增强其在痴呆症诊断中应用能力的研究方向。