Artificial intelligence-based methods have generated substantial interest in nuclear medicine. An area of significant interest has been using deep-learning (DL)-based approaches for denoising images acquired with lower doses, shorter acquisition times, or both. Objective evaluation of these approaches is essential for clinical application. DL-based approaches for denoising nuclear-medicine images have typically been evaluated using fidelity-based figures of merit (FoMs) such as RMSE and SSIM. However, these images are acquired for clinical tasks and thus should be evaluated based on their performance in these tasks. Our objectives were to (1) investigate whether evaluation with these FoMs is consistent with objective clinical-task-based evaluation; (2) provide a theoretical analysis for determining the impact of denoising on signal-detection tasks; (3) demonstrate the utility of virtual clinical trials (VCTs) to evaluate DL-based methods. A VCT to evaluate a DL-based method for denoising myocardial perfusion SPECT (MPS) images was conducted. The impact of DL-based denoising was evaluated using fidelity-based FoMs and AUC, which quantified performance on detecting perfusion defects in MPS images as obtained using a model observer with anthropomorphic channels. Based on fidelity-based FoMs, denoising using the considered DL-based method led to significantly superior performance. However, based on ROC analysis, denoising did not improve, and in fact, often degraded detection-task performance. The results motivate the need for objective task-based evaluation of DL-based denoising approaches. Further, this study shows how VCTs provide a mechanism to conduct such evaluations using VCTs. Finally, our theoretical treatment reveals insights into the reasons for the limited performance of the denoising approach.
翻译:基于人工智能的方法已在核医学领域引起广泛关注。其中,利用深度学习(DL)方法对低剂量、短采集时间或两者兼有的图像进行降噪处理已成为重要研究方向。这些方法的客观评估对临床应用至关重要。核医学图像降噪的深度学习方法通常使用基于保真度的品质因数(FoM)如均方根误差(RMSE)和结构相似性指数(SSIM)进行评估。然而,这些图像是为临床任务获取的,因此应基于其在任务中的表现进行评估。本研究旨在:(1)探究基于保真度的评估与基于目标的临床任务评估是否一致;(2)提供理论分析以确定降噪对信号检测任务的影响;(3)论证虚拟临床试验(VCT)在评估深度学习方法中的效用。我们构建了VCT框架,用于评估基于深度学习的心肌灌注SPECT(MPS)图像降噪方法。通过基于保真度的FoM和采用带拟人通道模型观察者获得的检测MPS图像灌注缺损性能的AUC,评估了深度学习降噪的影响。基于保真度的FoM结果显示,所考虑的深度学习方法实现了显著更优的降噪性能。然而,ROC分析表明,降噪并未改善,反而常降低检测任务性能。该结果揭示了基于目标的深度学习降噪方法客观评估的必要性。此外,本研究展示了VCT如何为开展此类评估提供有效机制。最后,理论分析揭示了该降噪方法性能受限的内在原因。