Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable patch-level reasoning remains challenging. End-to-end pathology MLLMs often hallucinate morphological features, while recent agentic systems usually merge tool outputs and retrieved knowledge into a shared context, making decisions vulnerable to conflicting evidence and context contamination. We propose PathoSage, a three-stage framework that explicitly separates knowledge retrieval, evidence collection, and evidence adjudication for patch-level pathology multimodal reasoning. Its core component, Structured Evidence Deliberation, independently evaluates heterogeneous evidence from tools, performs conflict analysis, and generates the final judgment in a fresh context to reduce anchoring bias. We further introduce a training-free Beta-Bernoulli experience system with continuous credit assignment to model long-term tool reliability and construct similarity-weighted priors for future tool use. Experiments show that PathoSage effectively mitigates VQA hallucinations and classifier disagreement, outperforming strong pathology MLLM and agentic baselines. Our results highlight explicit evidence adjudication and reliability-aware tool modeling as key ingredients for robust pathology agents.
翻译:多模态大语言模型与代理工作流的最新进展在计算病理学中展现出巨大潜力,但可靠的斑块级推理仍具挑战性。端到端病理学MLLM常出现形态特征幻觉,而现有代理系统通常将工具输出与检索知识合并到共享上下文中,导致决策易受矛盾证据与上下文污染影响。我们提出PathoSage——一种面向斑块级病理学多模态推理的三阶段框架,显式分离知识检索、证据收集与证据裁决。其核心组件"结构化证据审议"可独立评估来自工具的异构证据,执行冲突分析,并在全新上下文中生成最终判决以减轻锚定偏差。我们进一步提出免训练的Beta-Bernoulli经验系统,通过连续信用分配建模长期工具可靠性,并为未来工具使用构建相似性加权先验。实验表明,PathoSage能有效缓解VQA幻觉与分类器分歧,性能超越强病理学MLLM及代理基线。研究结果揭示了显式证据裁决与可靠性感知工具建模是构建稳健病理学代理的关键要素。