Large Language Models (LLMs) have demonstrated strong capabilities in biomedical question answering, yet their tendency to generate plausible but unverified claims poses serious risks in clinical settings. To mitigate these risks, the TREC 2025 BioGen track mandates grounded answers that explicitly surface contradictory evidence (Task A) and the generation of narrative driven, fully attributed responses (Task B). Addressing the absence of target ground truth, we present a proxy-based development framework using the SciFact dataset to systematically optimize retrieval architectures. Our iterative evaluation revealed a "Simplicity Paradox": complex adversarial dense retrieval strategies failed catastrophically at contradiction detection (MRR 0.023) due to Semantic Collapse, where negation signals become indistinguishable in vector space. We further identify a Retrieval Asymmetry: filtering dense embeddings improves contradiction detection but degrades support recall, compromising reliability. We resolve this via a Decoupled Lexical Architecture built on a unified BM25 backbone, balancing semantic support recall (0.810) with precise contradiction surfacing (0.750). This approach achieves the highest Weighted MRR (0.790) on the proxy benchmark while remaining the only viable strategy for scaling to the 30 million document PubMed corpus. For answer generation, we introduce Narrative Aware Reranking and One-Shot In-Context Learning, improving citation coverage from 50% (zero-shot) to 100%. Official TREC results confirm our findings: our system ranks 2nd on Task A contradiction F1 and 3rd out of 50 runs on Task B citation coverage (98.77%), achieving zero citation contradict rate. Our work transforms LLMs from stochastic generators into honest evidence synthesizers, showing that epistemic integrity in biomedical AI requires precision and architectural scalability isolated metric optimization.


翻译:大型语言模型(LLM)在生物医学问答中展现出强大能力,但其生成看似合理却未经核实的论断的倾向在临床应用中构成严重风险。为缓解这些风险,TREC 2025 BioGen赛道明确要求基于证据的答案(任务A:显式呈现矛盾证据)及生成叙事驱动、完全归因的响应(任务B)。针对目标标注缺失问题,我们提出基于SciFact数据集的代理开发框架,通过系统化优化检索架构。迭代评估揭示了"简单性悖论":复杂对抗性稠密检索策略在矛盾检测中完全失效(MRR 0.023),其根源在于语义坍缩——向量空间中否定信号变得难以区分。我们进一步识别出检索非对称性:过滤稠密嵌入可提升矛盾检测性能,但会降低支持证据召回率,从而削弱系统可靠性。为此,我们提出基于统一BM25骨干网络的解耦词汇架构,在语义支持证据召回率(0.810)与精准矛盾呈现(0.750)间实现平衡。该方案在代理基准上取得最高加权MRR(0.790),且是唯一可扩展至PubMed三千万文献语料的可行策略。在答案生成方面,我们采用叙事感知重排序与单样本上下文学习,将引文覆盖率从50%(零样本)提升至100%。TREC官方结果证实了我们的发现:系统在任务A矛盾F1值上排名第二,任务B引文覆盖率(98.77%)在50组运行中位列第三,并实现零引文矛盾率。本研究将LLM从随机生成器转化为诚实的证据综合器,表明生物医学AI的认知完整性既需精确性,也需架构可扩展性,而非孤立的指标优化。

0
下载
关闭预览

相关内容

具有动能的生命体。
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
大语言模型遇上知识图谱:问答系统中的融合与机遇
专知会员服务
30+阅读 · 2025年5月30日
定制化大型语言模型的图检索增强生成综述
专知会员服务
39+阅读 · 2025年1月28日
大型语言模型疾病诊断综述
专知会员服务
32+阅读 · 2024年9月21日
Nat. Med. | 医学中的大型语言模型
专知会员服务
58+阅读 · 2023年9月19日
深度学习模型不确定性方法对比
PaperWeekly
20+阅读 · 2020年2月10日
NLG ≠ 机器写作 | 专家专栏
量子位
13+阅读 · 2018年9月10日
国家自然科学基金
11+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
VIP会员
最新内容
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
2+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
11+阅读 · 7月22日
相关基金
国家自然科学基金
11+阅读 · 2016年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员