Retinal fundus images play a crucial role in the early detection of eye diseases and, using deep learning approaches, recent studies have even demonstrated their potential for detecting cardiovascular risk factors and neurological disorders. However, the impact of technical factors on these images can pose challenges for reliable AI applications in ophthalmology. For example, large fundus cohorts are often confounded by factors like camera type, image quality or illumination level, bearing the risk of learning shortcuts rather than the causal relationships behind the image generation process. Here, we introduce a novel population model for retinal fundus images that effectively disentangles patient attributes from camera effects, thus enabling controllable and highly realistic image generation. To achieve this, we propose a novel disentanglement loss based on distance correlation. Through qualitative and quantitative analyses, we demonstrate the effectiveness of this novel loss function in disentangling the learned subspaces. Our results show that our model provides a new perspective on the complex relationship between patient attributes and technical confounders in retinal fundus image generation.
翻译:眼底视网膜图像在眼病的早期检测中发挥着关键作用,近年来,基于深度学习方法的研究甚至已证明其在检测心血管风险因素和神经系统疾病方面的潜力。然而,技术因素对这些图像的影响可能为眼科领域的可靠人工智能应用带来挑战。例如,大规模眼底数据集常因相机类型、图像质量或光照水平等因素产生混杂效应,导致模型倾向于学习捷径而非图像生成过程背后的因果关系。本文提出了一种新颖的视网膜眼底图像群体模型,该模型能够有效解耦患者属性与相机效应,从而实现可控且高度逼真的图像生成。为此,我们基于距离相关性提出了一种新的解耦损失函数。通过定性和定量分析,我们证明了该新型损失函数在解耦学习到的子空间方面的有效性。结果表明,我们的模型为理解视网膜眼底图像生成中患者属性与技术混杂因素之间的复杂关系提供了全新视角。