Optic nerve head (ONH) detection has been a crucial area of study in ophthalmology for years. However, the significant discrepancy between fundus image datasets, each generated using a single type of fundus camera, poses challenges to the generalizability of ONH detection approaches developed based on semantic segmentation networks. Despite the numerous recent advancements in general-purpose semantic segmentation methods using convolutional neural networks (CNNs) and Transformers, there is currently a lack of benchmarks for these state-of-the-art (SoTA) networks specifically trained for ONH detection. Therefore, in this article, we make contributions from three key aspects: network design, the publication of a dataset, and the establishment of a comprehensive benchmark. Our newly developed ONH detection network, referred to as ODFormer, is based upon the Swin Transformer architecture and incorporates two novel components: a multi-scale context aggregator and a lightweight bidirectional feature recalibrator. Our published large-scale dataset, known as TongjiU-DROD, provides multi-resolution fundus images for each participant, captured using two distinct types of cameras. Our established benchmark involves three datasets: DRIONS-DB, DRISHTI-GS1, and TongjiU-DROD, created by researchers from different countries and containing fundus images captured from participants of diverse races and ages. Extensive experimental results demonstrate that our proposed ODFormer outperforms other state-of-the-art (SoTA) networks in terms of performance and generalizability. Our dataset and source code are publicly available at mias.group/ODFormer.
翻译:视神经盘(ONH)检测多年来一直是眼科学研究的关键领域。然而,由于不同眼底图像数据集均使用单一类型的眼底相机采集,数据集之间存在显著差异,这对基于语义分割网络开发的ONH检测方法的泛化能力提出了挑战。尽管近期基于卷积神经网络(CNN)和Transformer的通用语义分割方法取得了诸多进展,但目前仍缺乏专门为ONH检测训练的这些最先进(SoTA)网络的基准测试。因此,本文从三个关键方面做出贡献:网络设计、数据集发布以及综合基准的建立。我们新开发的ONH检测网络(称为ODFormer)基于Swin Transformer架构,并包含两个新颖组件:多尺度上下文聚合器和轻量级双向特征重校准器。我们发布的大规模数据集(称为TongjiU-DROD)为每位参与者提供了多分辨率眼底图像,这些图像使用两种不同类型的相机采集。我们建立的基准包含三个数据集:DRIONS-DB、DRISHTI-GS1和TongjiU-DROD,这些数据集由不同国家的研究人员创建,包含从不同种族和年龄参与者采集的眼底图像。大量实验结果表明,我们提出的ODFormer在性能和泛化能力方面均优于其他最先进(SoTA)网络。我们的数据集和源代码已在mias.group/ODFormer公开提供。