In high-stakes information domains such as healthcare, where large language models (LLMs) can produce hallucinations or misinformation, retrieval-augmented generation (RAG) has been proposed as a mitigation strategy, grounding model outputs in external, domain-specific documents. Yet, this approach can introduce errors when source documents contain outdated or contradictory information. This work investigates the performance of five LLMs in generating RAG-based responses to medicine-related queries. Our contributions are three-fold: i) the creation of a benchmark dataset using consumer medicine information documents from the Australian Therapeutic Goods Administration (TGA), where headings are repurposed as natural language questions, ii) the retrieval of PubMed abstracts using TGA headings, stratified across multiple publication years, to enable controlled temporal evaluation of outdated evidence, and iii) a comparative analysis of the frequency and impact of outdated or contradictory content on model-generated responses, assessing how LLMs integrate and reconcile temporally inconsistent information. Our findings show that contradictions between highly similar abstracts do, in fact, degrade performance, leading to inconsistencies and reduced factual accuracy in model answers. These results highlight that retrieval similarity alone is insufficient for reliable medical RAG and underscore the need for contradiction-aware filtering strategies to ensure trustworthy responses in high-stakes domains.
翻译:在高风险信息领域(如医疗健康)中,大语言模型可能产生幻觉或错误信息,因此检索增强生成被提出作为缓解策略——通过将模型输出锚定在外部领域特定文档中。然而,当源文档包含过时或矛盾信息时,这种方法可能引入错误。本研究探究了五种大语言模型在生成基于检索增强的医学查询响应时的表现。我们的贡献包含三方面:第一,利用澳大利亚治疗用品管理局的消费者药品信息文档创建基准数据集,将标题重构为自然语言问题;第二,使用治疗用品管理局标题检索PubMed摘要,按多个发表年份分层,以实现对过时证据的受控时间评估;第三,对过时或矛盾内容在模型生成响应中的出现频率及影响进行对比分析,评估大语言模型整合与调和时序不一致信息的能力。实验发现:高度相似摘要间的矛盾确实会降低模型性能,导致答案不一致和事实准确性下降。这些结果表明,仅依靠检索相似度不足以支持可靠的医疗检索增强生成,并凸显了在高风险领域需要采用矛盾感知过滤策略以确保可信响应。