Diabetic Retinopathy (DR) is an aggressive retinal disease and a leading cause of global blindness, yet its clinical management is currently hindered by the black-box nature of diagnostic AI. While deep learning models achieve high classification accuracy, there is a critical lack of explainability methods capable of detailing the exact anatomical landmarks and lesion distributions that lead to a clinical decision for DR. Therefore, we propose HSQ-VLM, a novel quadrant segmentation pipeline on fundus images that utilizes a Landmark-Anchored Cartesian Cross-Attention mechanism to unify visual feature extraction with structured clinical reasoning. Unlike traditional methods that rely on arbitrary image partitioning, our pipeline implements 4-quadrant Topological Latent Partitioning (TLP) to dynamically align retinal features with a fovea-centered coordinate system. This allows the Vision-Language Model to generate natural language reports that quantify pathology with anatomical precision. On a dataset of 3,500 high-resolution fundus images, this innovative methodology achieved a lesion detection sensitivity of 99.6% for hemorrhages and 96.4% for microaneurysms, while demonstrating a significant reduction in boundary-ambiguity errors compared to standard segmentation baselines.
翻译:糖尿病视网膜病变是一种侵袭性视网膜疾病,也是全球失明的主要原因,然而其临床管理目前受到诊断AI黑箱性质的阻碍。尽管深度学习模型实现了高分类准确率,但在能够详细说明导致DR临床决策的精确解剖标志和病变分布的可解释性方法上仍存在严重缺失。因此,我们提出HSQ-VLM——一种基于眼底图像的新型象限分割流水线,该流水线利用地标锚定笛卡尔交叉注意力机制,将视觉特征提取与结构化临床推理相统一。与传统依赖任意图像分区的方法不同,我们的流水线采用四象限拓扑潜在分割(TLP),动态地将视网膜特征与以中心凹为中心的坐标系对齐。这使得视觉语言模型能够生成量化病理并具备解剖学精确性的自然语言报告。在包含3,500张高分辨率眼底图像的数据集上,该创新方法对出血点和微动脉瘤的病变检测敏感度分别达到99.6%和96.4%,同时与标准分割基线相比,边界模糊误差显著降低。