Automated medical report generation for 3D PET/CT imaging is fundamentally challenged by the high-dimensional nature of volumetric data and a critical scarcity of annotated datasets, particularly for low-resource languages. Current black-box methods map whole volumes to reports, ignoring the clinical workflow of analyzing localized Regions of Interest (RoIs) to derive diagnostic conclusions. In this paper, we bridge this gap by introducing VietPET-RoI, the first large-scale 3D PET/CT dataset with fine-grained RoI annotation for a low-resource language, comprising 600 PET/CT samples and 1,960 manually annotated RoIs, paired with corresponding clinical reports. Furthermore, to demonstrate the utility of this dataset, we propose HiRRA, a novel framework that mimics the professional radiologist diagnostic workflow by employing graph-based relational modules to capture dependencies between RoI attributes. This approach shifts from global pattern matching toward localized clinical findings. Additionally, we introduce new clinical evaluation metrics, namely RoI Coverage and RoI Quality Index, that measure both RoI localization accuracy and attribute description fidelity using LLM-based extraction. Extensive evaluation demonstrates that our framework achieves SOTA performance, surpassing existing models by 19.7% in BLEU and 4.7% in ROUGE-L, while achieving a remarkable 45.8% improvement in clinical metrics, indicating enhanced clinical reliability and reduced hallucination. Our code and dataset are available on GitHub.
翻译:三维PET/CT影像的自动化医学报告生成面临根本性挑战,主要源于体数据的高维特性以及标注数据集(尤其是低资源语言环境)的严重匮乏。现有黑盒方法直接将完整影像映射为报告,忽略了临床工作流程中分析局部感兴趣区域(RoI)以得出诊断结论的实践。本文通过引入VietPET-RoI数据集填补这一空白——这是首个面向低资源语言、包含细粒度RoI标注的大规模三维PET/CT数据集,涵盖600个PET/CT样本与1,960个手工标注的RoI,并配有对应临床报告。为展示该数据集的实用价值,我们提出HiRRA框架,该框架通过采用基于图的关联模块捕捉RoI属性间的依赖关系,模拟专业放射科医师的诊断工作流程。该方法将全局模式匹配转向局部化临床发现。此外,我们引入全新临床评估指标——RoI覆盖率与RoI质量指数,利用基于大语言模型(LLM)的提取技术,同时度量RoI定位精度与属性描述保真度。大量评估表明,本框架实现了最先进性能,在BLEU和ROUGE-L指标上分别超越现有模型19.7%和4.7%,临床指标提升幅度高达45.8%,展现了增强的临床可靠性与显著减少的幻觉现象。我们的代码和数据集已在GitHub上公开。