Automatic medical report generation (MRG) is of great research value as it has the potential to relieve radiologists from the heavy burden of report writing. Despite recent advancements, accurate MRG remains challenging due to the need for precise clinical understanding and the identification of clinical findings. Moreover, the imbalanced distribution of diseases makes the challenge even more pronounced, as rare diseases are underrepresented in training data, making their diagnostic performance unreliable. To address these challenges, we propose diagnosis-driven prompts for medical report generation (PromptMRG), a novel framework that aims to improve the diagnostic accuracy of MRG with the guidance of diagnosis-aware prompts. Specifically, PromptMRG is based on encoder-decoder architecture with an extra disease classification branch. When generating reports, the diagnostic results from the classification branch are converted into token prompts to explicitly guide the generation process. To further improve the diagnostic accuracy, we design cross-modal feature enhancement, which retrieves similar reports from the database to assist the diagnosis of a query image by leveraging the knowledge from a pre-trained CLIP. Moreover, the disease imbalanced issue is addressed by applying an adaptive logit-adjusted loss to the classification branch based on the individual learning status of each disease, which overcomes the barrier of text decoder's inability to manipulate disease distributions. Experiments on two MRG benchmarks show the effectiveness of the proposed method, where it obtains state-of-the-art clinical efficacy performance on both datasets.
翻译:自动医学报告生成(MRG)具有重要的研究价值,因为它有潜力减轻放射科医生撰写报告的沉重负担。尽管近年来取得了进展,但由于需要精准的临床理解和临床发现的识别,准确的MRG仍面临挑战。此外,疾病的分布不平衡使这一挑战更加显著,因为罕见疾病在训练数据中代表性不足,导致其诊断性能不可靠。为了解决这些问题,我们提出了诊断驱动提示的医学报告生成(PromptMRG),这是一个新颖的框架,旨在通过诊断感知提示的引导提升MRG的诊断准确性。具体来说,PromptMRG基于编码器-解码器架构,并增加了一个额外的疾病分类分支。在生成报告时,分类分支的诊断结果被转换为令牌提示,以明确指导生成过程。为进一步提升诊断准确性,我们设计了跨模态特征增强机制,从数据库中检索相似报告,利用预训练CLIP的知识辅助查询图像的诊断。此外,通过针对每种疾病的个体学习状态,在分类分支上应用自适应对数调整损失,解决了疾病不平衡问题,克服了文本解码器无法操控疾病分布的障碍。在两个MRG基准上的实验表明,所提方法的有效性,在两个数据集上均获得了最先进的临床疗效性能。