Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, which comprises four subtasks: question interpretation, evidence identification, answer generation, and evidence alignment. Our approach decouples the task into independent, modular stages and employs DSPy"s MIPROv2 optimizer to automatically discover high-performing prompts, jointly tuning instructions and few-shot demonstrations for each stage. Within every stage, self-consistency voting over multiple stochastic inference runs suppresses spurious errors and improves reliability, while stage-specific verification mechanisms (e.g., self-reflection and chain-of-verification for alignment) further refine output quality. Among all teams that participated in all four subtasks, our method ranks second overall (mean rank 4.00), placing 4th, 1st, 4th, and 7th on Subtasks 1-4, respectively. These results demonstrate that systematic, per-stage prompt optimization combined with self-consistency mechanisms is a cost-effective alternative to model fine-tuning for multifaceted clinical QA.
翻译:面向电子健康记录的自动问答要求精准的证据检索、忠实的答案生成以及将答案明确锚定在临床记录中。本文介绍了我们为CL4Health@LREC 2026的ArchEHR-QA 2026共享任务提出的Neural1.5方法,该任务包含四个子任务:问题解释、证据识别、答案生成和证据对齐。我们的方法将任务解耦为独立、模块化的阶段,采用DSPy的MIPROv2优化器自动发现高性能提示,联合调整每个阶段的指令和少样本示例。在每个阶段中,基于多次随机推理的自一致性投票可抑制虚假错误并提升可靠性,而阶段特定的验证机制(如自反思和对齐的链式验证)进一步优化输出质量。在所有参与四个子任务的队伍中,我们的方法总体排名第二(平均排名4.00),在子任务1-4上分别位列第4、第1、第4和第7名。这些结果表明,对于多方面的临床问答,系统化分阶段提示优化结合自一致性机制是模型微调的一种低成本高效替代方案。