Diabetic Macular Edema (DME) is a leading cause of vision loss among patients with Diabetic Retinopathy (DR). While deep learning has shown promising results for automatically detecting this condition from fundus images, its application remains challenging due the limited availability of annotated data. Foundation Models (FM) have emerged as an alternative solution. However, it is unclear if they can cope with DME detection in particular. In this paper, we systematically compare different FM and standard transfer learning approaches for this task. Specifically, we compare the two most popular FM for retinal images-RETFound and FLAIR-and an EfficientNetB0 backbone, across different training regimes and evaluation settings in IDRiD, MESSIDOR-2 and OCT-and-Eye-FundusImages (OEFI). Results show that despite their scale, FM do not consistently outperform fine-tuned CNNs in this task. In particular, EfficientNet-B0 consistently achieves competitive or superior performance across evaluation settings, with FLAIR being the most competitive foundation model, consistently outperforming RETFound. These findings suggest that FMs do not necessarily provide an advantage for fine-grained ophthalmic tasks such as DME detection, even after fine-tuning, highlighting lightweight CNNs as strong baselines in data-scarce environments.
翻译:暂无翻译