Automated 3D radiology report generation often suffers from clinical hallucinations and a lack of the iterative verification found in human practice. While recent Vision-Language Models (VLMs) have advanced the field, they typically operate as monolithic "black-box" systems without the collaborative oversight characteristic of clinical workflows. To address these challenges, we propose MARCH (Multi-Agent Radiology Clinical Hierarchy), a multi-agent framework that emulates the professional hierarchy of radiology departments and assigns specialized roles to distinct agents. MARCH utilizes a Resident Agent for initial drafting with multi-scale CT feature extraction, multiple Fellow Agents for retrieval-augmented revision, and an Attending Agent that orchestrates an iterative, stance-based consensus discourse to resolve diagnostic discrepancies. On the RadGenome-ChestCT dataset, MARCH significantly outperforms state-of-the-art baselines in both clinical fidelity and linguistic accuracy. Our work demonstrates that modeling human-like organizational structures enhances the reliability of AI in high-stakes medical domains.
翻译:摘要:自动化三维放射学报告生成常受临床幻觉困扰,且缺乏人类实践中的迭代验证机制。尽管近期视觉语言模型(VLMs)推动了该领域发展,但其通常以单一“黑箱”系统运行,缺乏临床工作流程中特有的协作监督。为应对这些挑战,我们提出MARCH(多智能体放射科临床层次架构),这是一种模拟放射科室专业层级的多智能体框架,为不同智能体分配专业化角色。MARCH利用住院医师智能体执行基于多尺度CT特征提取的初始草稿撰写,多位进修医师智能体负责检索增强式修订,并由主治医师智能体协调基于立场共识的迭代式论述,以解决诊断分歧。在RadGenome-ChestCT数据集上,MARCH在临床保真度和语言准确性两方面均显著超越当前最优基线模型。我们的研究表明,模拟人类组织架构能提升高风险医疗领域中AI系统的可靠性。