In this work, we propose Oph-Guid-RAG, a multimodal visual RAG system for ophthalmology clinical question answering and decision support. We treat each guideline page as an independent evidence unit and directly retrieve page images, preserving tables, flowcharts, and layout information. We further design a controllable retrieval framework with routing and filtering, which selectively introduces external evidence and reduces noise. The system integrates query decomposition, query rewriting, retrieval, reranking, and multimodal reasoning, and provides traceable outputs with guideline page references. We evaluate our method on HealthBench using a doctor-based scoring protocol. On the hard subset, our approach improves the overall score from 0.2969 to 0.3861 (+0.0892, +30.0%) compared to GPT-5.2, and achieves higher accuracy, improving from 0.5956 to 0.6576 (+0.0620, +10.4%). Compared to GPT-5.4, our method achieves a larger accuracy gain of +0.1289 (+24.4%). These results show that our method is more effective on challenging cases that require precise, evidence-based reasoning. Ablation studies further show that reranking, routing, and retrieval design are critical for stable performance, especially under difficult settings. Overall, we show how combining visionbased retrieval with controllable reasoning can improve evidence grounding and robustness in clinical AI applications,while pointing out that further work is needed to be more complete.
翻译:本文提出Oph-Guid-RAG,一种面向眼科临床问答与决策支持的多模态视觉检索增强生成(RAG)系统。我们将每页临床指南作为独立证据单元,直接检索页面图像,保留表格、流程图及布局信息。进一步设计具有路由与过滤机制的可控检索框架,选择性引入外部证据并降低噪声。系统整合查询分解、查询改写、检索、重排序与多模态推理,并输出附带指南页面引用的可追溯结果。我们采用基于医生评分的协议在HealthBench上评估该方法。在困难子集上,相较于GPT-5.2,我们的方法将整体评分从0.2969提升至0.3861(+0.0892,+30.0%),准确率从0.5956提升至0.6576(+0.0620,+10.4%)。与GPT-5.4相比,准确率提升幅度更大(+0.1289,+24.4%)。结果表明,在需要精准循证推理的挑战性病例上,我们的方法更具优势。消融实验进一步表明,重排序、路由与检索设计对稳定性能至关重要,尤其在困难场景下。综上,我们展示了将基于视觉的检索与可控推理相结合可提升临床AI应用的证据归因能力与鲁棒性,同时指出需进一步研究以实现更完备的系统。