Recent advances in large multimodal models have allowed for the development of interactive chat models that can converse and reason about pathology whole-slide images (WSIs). However, existing slide-level chat systems are often highly specialized, typically compressing WSIs into fixed slide-level embeddings or relying on multi-component pipelines, which can lose multi-scale detail and limit generalizability beyond the target task. We present GIANT (Gigapixel Image Agent for Navigating Tissue), a simple, training-free approach that lets general-purpose multimodal models navigate WSIs on their own, iteratively selecting multi-magnification crops and aggregating evidence over time. To evaluate generalizability in WSI question answering and to promote reproducibility, we introduce MultiPathQA, a benchmark suite spanning five clinical challenges and 934 questions over 868 unique WSIs. This includes a new set of 128 pathologist-authored multiple-choice questions designed to mirror real diagnostic search and multi-scale reasoning. Using GPT-5, GIANT outperforms models specialized for pathology question answering, achieving state-of-the-art performance on four out of five benchmarks.
翻译:大规模多模态模型的近期进展推动了交互式聊天模型的发展,使其能够就病理学全切片图像(WSI)进行对话与推理。然而,现有切片级聊天系统往往高度专业化,通常将全切片图像压缩为固定切片级嵌入特征,或依赖多组件流水线架构,这可能导致多尺度细节丢失,并限制其在目标任务之外的泛化能力。本文提出GIANT(组织导航千亿像素图像智能体),一种无需训练的简洁方法,使通用多模态模型能自主导航全切片图像,通过迭代式选择多放大倍数图像块并随时间累积证据。为评估全切片图像问答的泛化能力并促进可复现性,我们推出MultiPathQA基准套件,涵盖五大临床挑战场景、934个问题及868张独立全切片图像,其中包含128道由病理学家编写的多选题,旨在模拟真实诊断搜索与多尺度推理过程。基于GPT-5模型,GIANT在病理学问答专用模型中表现优异,在五项基准测试中四项达到最先进水平。