Data scarcity and privacy concerns limit the availability of high-quality medical images for public use, which can be mitigated through medical image synthesis. However, current medical image synthesis methods often struggle to accurately capture the complexity of detailed anatomical structures and pathological conditions. To address these challenges, we propose a novel medical image synthesis model that leverages fine-grained image-text alignment and anatomy-pathology prompts to generate highly detailed and accurate synthetic medical images. Our method integrates advanced natural language processing techniques with image generative modeling, enabling precise alignment between descriptive text prompts and the synthesized images' anatomical and pathological details. The proposed approach consists of two key components: an anatomy-pathology prompting module and a fine-grained alignment-based synthesis module. The anatomy-pathology prompting module automatically generates descriptive prompts for high-quality medical images. To further synthesize high-quality medical images from the generated prompts, the fine-grained alignment-based synthesis module pre-defines a visual codebook for the radiology dataset and performs fine-grained alignment between the codebook and generated prompts to obtain key patches as visual clues, facilitating accurate image synthesis. We validate the superiority of our method through experiments on public chest X-ray datasets and demonstrate that our synthetic images preserve accurate semantic information, making them valuable for various medical applications.
翻译:数据稀缺和隐私问题限制了高质量医学图像的公开可用性,而医学图像合成可缓解这一困境。然而,当前医学图像合成方法往往难以精确捕捉精细解剖结构及病理状态的复杂性。为应对这些挑战,我们提出一种新型医学图像合成模型,通过细粒度图文对齐与解剖-病理提示生成高保真合成医学图像。该方法融合先进自然语言处理技术与图像生成建模,实现描述性文本提示与合成图像解剖及病理细节间的精准对齐。模型包含两个核心组件:解剖-病理提示模块与基于细粒度对齐的合成模块。前者自动为高质量医学图像生成描述性提示,后者在放射学数据集上预定义视觉码本,并通过细粒度对齐将码本与生成提示关联,提取关键图像块作为视觉线索,最终合成高质量医学图像。在公开胸部X光数据集上的实验验证了本方法的优越性,结果表明所合成的图像保留了精确的语义信息,可广泛应用于各类医学任务。