Despite the rapid progress in self-supervised learning (SSL), end-to-end fine-tuning still remains the dominant fine-tuning strategy for medical imaging analysis. However, it remains unclear whether this approach is truly optimal for effectively utilizing the pre-trained knowledge, especially considering the diverse categories of SSL that capture different types of features. In this paper, we first establish strong contrastive and restorative SSL baselines that outperform SOTA methods across four diverse downstream tasks. Building upon these strong baselines, we conduct an extensive fine-tuning analysis across multiple pre-training and fine-tuning datasets, as well as various fine-tuning dataset sizes. Contrary to the conventional wisdom of fine-tuning only the last few layers of a pre-trained network, we show that fine-tuning intermediate layers is more effective, with fine-tuning the second quarter (25-50%) of the network being optimal for contrastive SSL whereas fine-tuning the third quarter (50-75%) of the network being optimal for restorative SSL. Compared to the de-facto standard of end-to-end fine-tuning, our best fine-tuning strategy, which fine-tunes a shallower network consisting of the first three quarters (0-75%) of the pre-trained network, yields improvements of as much as 5.48%. Additionally, using these insights, we propose a simple yet effective method to leverage the complementary strengths of multiple SSL models, resulting in enhancements of up to 3.57% compared to using the best model alone. Hence, our fine-tuning strategies not only enhance the performance of individual SSL models, but also enable effective utilization of the complementary strengths offered by multiple SSL models, leading to significant improvements in self-supervised medical imaging analysis.
翻译:尽管自监督学习(SSL)取得了快速进展,端到端微调仍是医学影像分析中占主导地位的微调策略。然而,这种方法是否真正能最优地利用预训练知识仍不明确,特别是考虑到不同类型的SSL会捕获不同特征。本文首先建立了强大的对比式和修复式SSL基线,在四个不同下游任务中均优于现有最优方法。基于这些强基线,我们在多个预训练和微调数据集以及不同规模的微调数据集上进行了广泛的微调分析。与传统观念中仅微调预训练网络最后几层不同,我们证明微调中间层更为有效:对于对比式SSL,微调网络第二季度(25-50%)最优;而对于修复式SSL,微调第三季度(50-75%)最佳。与端到端微调这一事实标准相比,我们提出的最优微调策略——即仅微调由预训练网络前四分之三(0-75%)构成的浅层网络——性能提升高达5.48%。此外,基于这些发现,我们提出一种简单有效的方法来利用多个SSL模型的互补优势,与单独使用最优模型相比,性能提升达3.57%。因此,我们的微调策略不仅增强了单个SSL模型的性能,还能有效利用多个SSL模型提供的互补优势,从而显著提升自监督医学影像分析的效果。