Generative models for volumetric medical images have found many applications in medical imaging, ranging from data augmentation to serving as priors for inverse problems. For these applications, generating high-resolution 3D images with strong controllability is essential but remains highly challenging. Existing approaches typically control generation either through radiology reports used as text prompts or through full image segmentation. While text-based prompting is flexible, it provides limited spatial control over the location, shape, and boundary of abnormalities. In contrast, segmentation-based methods receive precise spatial guidance but are restrictive in requiring full-organ annotations. In this work, we propose a flexible multimodal framework for controllable volumetric image generation that supports input from radiology reports and segmentation prompts (both optional). Our approach allows users to provide segmentation of a specific anatomy or abnormality without requiring full-organ annotations. The semantic meaning of the segmentation mask is specified through an accompanying text description, resulting in a highly flexible and scalable conditioning mechanism. We develop a memory-efficient architecture based on a modified diffusion transformer that jointly processes image and segmentation tokens. The model further incorporates gated attention to effectively attend to long radiology reports. Experiments demonstrate that our method achieves state-of-the-art perceptual and semantic scores (e.g., 24% relative improvement in mean FID), generates high-resolution anatomically consistent CT volumes, and improves data efficiency when used for data augmentation. Radiologists' evaluation further confirms strong alignment between generated and real medical images.


翻译:医学影像生成模型在数据增强、逆问题先验等医学成像领域具有广泛应用。为实现这些应用,生成高分辨率且强可控性的三维图像至关重要,但仍极具挑战性。现有方法通常通过放射学报告文本提示或完整图像分割来控制生成过程。基于文本的提示虽灵活,但难以对病灶的位置、形状和边界提供精确空间控制;而基于分割的方法虽具备精确空间引导,却受限于需要全器官标注。本文提出一种灵活的多模态框架,支持通过放射学报告和分割提示(均为可选输入)实现可控体积图像生成。该框架允许用户仅标注特定解剖结构或病灶区域,无需完整器官分割。分割掩膜的语义信息通过配套文本描述指定,从而构建高度灵活且可扩展的条件机制。我们基于改进的扩散Transformer架构开发了内存高效模型,可联合处理图像与分割令牌,并通过门控注意力机制有效关注长文本放射学报告。实验表明,该方法在感知与语义评分上达到最优水平(如平均FID相对提升24%),生成高分辨率解剖一致CT影像,并在数据增强场景下提升数据利用效率。放射科医师评估进一步证实生成影像与真实医学影像具有高度一致性。

0
下载
关闭预览

相关内容

【ETHZ博士论文】真实世界约束下的2D和3D生成模型
专知会员服务
25+阅读 · 2024年9月2日
【MIT博士论文】利用深度学习改进医学影像分割,165页pdf
专知会员服务
51+阅读 · 2021年8月28日
专知会员服务
117+阅读 · 2021年1月11日
超像素、语义分割、实例分割、全景分割 傻傻分不清?
计算机视觉life
19+阅读 · 2018年11月27日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关资讯
超像素、语义分割、实例分割、全景分割 傻傻分不清?
计算机视觉life
19+阅读 · 2018年11月27日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员