Most medical AI systems improve by scaling additional machinery: more fine-tuning data, more agents, and/or larger retrieval databases. In rare-disease diagnosis, however, such scaling can produce systems that are difficult to deploy, audit, and maintain. We asked whether state-of-the-art diagnostic performance could instead be achieved by extending the reasoning chain of a single AI agent: guiding it with a diagnostic policy, developed through human-AI collaboration and augmenting with freely available biomedical tools. We introduce LiteOdyssey, a lightweight rare-disease diagnostic framework that guides reasoning language model through a clinical genetics workflow. This framework was developed through Policy Iteration with Human Feedback (PIHF) and uses dynamic access to public biomedical tools. On two challenging benchmarks that provide only patient clinical features, LiteOdyssey achieved state-of-the-art performance, with an overall disease Recall@1 of 59.3% over the combined 1,243 cases of LIRICAL (n = 370) and the PhenoPacket Store (n = 873). Both benchmarks have a high proportion of ultra-rare disease (a prevalence below 1 in 1,000,000, with ultra-rare shares of approximately 45% and 52.8%, respectively). On the more difficult PhenoPacket subset, where causal diseases were not mapped to Orphanet in our rarity-mapping pipeline, LiteOdyssey achieved 60.7% Recall@1, compared with 10.7% for the same baseline model (GPT-5.4) without tools. This performance was achieved without fine-tuning, multi-agent ensembles, or a large case-retrieval database. Gains were also observed in the following: on cases never seen during development, on a private cohort of real-world rare disease patients, and on a smaller open-weights model. LiteOdyssey suggests a path toward rare-disease AI systems that are accurate, easier to deploy, and more transparent for physician review.
翻译:摘要:大多数医疗AI系统通过扩展额外机制(更多微调数据、更多智能体及/或更大规模的检索数据库)来提升性能。然而在罕见病诊断中,此类扩展可能导致系统难以部署、审计和维护。我们探究是否可通过延长单一AI智能体的推理链来实现最先进的诊断性能:通过人机协作制定诊断策略进行引导,并辅以免费可用的生物医学工具。我们提出LiteOdyssey——一个轻量级罕见病诊断框架,通过临床遗传学工作流引导推理语言模型。该框架采用人类反馈策略迭代(PIHF)开发,并可动态访问公共生物医学工具。在仅提供患者临床特征的两项具有挑战性的基准测试中,LiteOdyssey实现了最先进的性能:在LIRICAL(n=370)与PhenoPacket Store(n=873)共1243例病例组成的综合数据集上,疾病Recall@1整体达59.3%。两项基准测试均包含高比例超罕见病(患病率低于百万分之一,超罕见占比分别约45%和52.8%)。在更具难度的PhenoPacket子集(我们的罕见性映射流程中因果疾病未映射至Orphanet)上,LiteOdyssey达到60.7%的Recall@1,而使用相同基线模型(GPT-5.4)但无工具辅助时仅为10.7%。该性能的实现无需微调、多智能体集成或大规模病例检索数据库。在开发过程中未涉及的病例、真实世界罕见病患者私有队列以及小型开源权重模型上均观察到性能提升。LiteOdyssey为构建精准、易部署且便于医生审查的罕见病AI系统指明了方向。