Training Artificial Intelligence (AI) models on 3D images presents unique challenges compared to the 2D case: Firstly, the demand for computational resources is significantly higher, and secondly, the availability of large datasets for pre-training is often limited, impeding training success. This study proposes a simple approach of adapting 2D networks with an intermediate feature representation for processing 3D images. Our method employs attention pooling to learn to assign each slice an importance weight and, by that, obtain a weighted average of all 2D slices. These weights directly quantify the contribution of each slice to the contribution and thus make the model prediction inspectable. We show on all 3D MedMNIST datasets as benchmark and two real-world datasets consisting of several hundred high-resolution CT or MRI scans that our approach performs on par with existing methods. Furthermore, we compare the in-built interpretability of our approach to HiResCam, a state-of-the-art retrospective interpretability approach.
翻译:相较于二维情况,在三维图像上训练人工智能(AI)模型面临独特挑战:首先,计算资源需求显著增加;其次,可用于预训练的大规模数据集往往有限,阻碍了训练成功。本研究提出一种通过中间特征表示适配二维网络以处理三维图像的简单方法。该方法采用注意力池化来学习为每个切片分配重要性权重,从而获得所有二维切片的加权平均值。这些权重直接量化了每个切片对最终预测的贡献,使模型预测结果具备可检查性。我们在所有3D MedMNIST基准数据集以及两个包含数百张高分辨率CT或MRI扫描的真实数据集上证明,该方法性能与现有方法持平。此外,我们将该方法的固有可解释性与HiResCam(最先进的回溯性可解释方法)进行了对比。