In this paper, we propose a novel approach (called GPT4MIA) that utilizes Generative Pre-trained Transformer (GPT) as a plug-and-play transductive inference tool for medical image analysis (MIA). We provide theoretical analysis on why a large pre-trained language model such as GPT-3 can be used as a plug-and-play transductive inference model for MIA. At the methodological level, we develop several technical treatments to improve the efficiency and effectiveness of GPT4MIA, including better prompt structure design, sample selection, and prompt ordering of representative samples/features. We present two concrete use cases (with workflow) of GPT4MIA: (1) detecting prediction errors and (2) improving prediction accuracy, working in conjecture with well-established vision-based models for image classification (e.g., ResNet). Experiments validate that our proposed method is effective for these two tasks. We further discuss the opportunities and challenges in utilizing Transformer-based large language models for broader MIA applications.
翻译:本文提出一种新颖方法(称为GPT4MIA),利用生成式预训练Transformer(GPT)作为即插即用的转导推理工具,应用于医学图像分析(MIA)。我们从理论上论证了为何像GPT-3这样的大型预训练语言模型可被用作MIA的即插即用转导推理模型。在方法论层面,我们开发了若干技术处理手段以提升GPT4MIA的效率与效能,包括优化提示结构设计、样本选取以及代表性样本/特征的提示排序。我们给出了GPT4MIA的两个具体使用案例(附工作流程):(1)检测预测错误;(2)与成熟的基于视觉的图像分类模型(如ResNet)协同工作,提升预测准确性。实验验证了所提方法在上述两项任务中的有效性。我们进一步探讨了基于Transformer的大型语言模型在更广泛MIA应用中的机遇与挑战。