Recently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology NLI task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4's reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) GPT-4 outperforms ChatGPT in the radiology NLI task; 2) other specifically fine-tuned models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings demonstrate that constructing a generic model that is capable of solving various tasks across different domains is feasible.
翻译:近期,ChatGPT和GPT-4因其在语言处理领域无与伦比的表现而涌现并引发全球广泛关注。尽管它们在各类开放领域任务中展现了令人印象深刻的能力,但在放射学这类高度专业化领域中的适用性仍未得到验证。放射学因其特异性和复杂性,呈现出区别于开放领域数据的独特语言现象。评估大语言模型在如此特定领域中的表现,不仅有助于全面评价其整体性能,还能为未来的模型设计方向提供重要启示:即模型设计应趋向通用化还是领域专用化。为此,本研究在放射学自然语言推理任务上评估了ChatGPT/GPT-4的性能,并将其与专门在任务相关数据样本上微调的其他模型进行比较。我们还通过引入不同推理难度等级,对ChatGPT/GPT-4的推理能力进行了全面探究。结果表明:1) GPT-4在放射学自然语言推理任务中优于ChatGPT;2) 其他专门微调模型需要大量数据样本才能达到与ChatGPT/GPT-4相当的性能。这些发现证明,构建能够解决跨领域各类任务的通用模型是可行的。