GenerateCT, the first approach to generating 3D medical imaging conditioned on free-form medical text prompts, incorporates a text encoder and three key components: a novel causal vision transformer for encoding 3D CT volumes, a text-image transformer for aligning CT and text tokens, and a text-conditional super-resolution diffusion model. Given the absence of directly comparable methods in 3D medical imaging, we established baselines with cutting-edge methods to demonstrate our method's effectiveness. GenerateCT significantly outperforms these methods across all key metrics. Importantly, we explored GenerateCT's clinical applications by evaluating its utility in a multi-abnormality classification task. First, we established a baseline by training a multi-abnormality classifier on our real dataset. To further assess the model's generalization to external datasets and its performance with unseen prompts in a zero-shot scenario, we employed an external dataset to train the classifier, setting an additional benchmark. We conducted two experiments in which we doubled the training datasets by synthesizing an equal number of volumes for each set using GenerateCT. The first experiment demonstrated an 11% improvement in the AP score when training the classifier jointly on real and generated volumes. The second experiment showed a 7% improvement when training on both real and generated volumes based on unseen prompts. Moreover, GenerateCT enables the scaling of synthetic training datasets to arbitrary sizes. As an example, we generated 100,000 3D CT volumes, fivefold the number in our real dataset, and trained the classifier exclusively on these synthetic volumes. Impressively, this classifier surpassed the performance of the one trained on all available real data by a margin of 8%. Lastly, domain experts evaluated the generated volumes, confirming a high degree of alignment with the text prompt.
翻译:GenerateCT是首个基于自由形式医学文本提示生成3D医学影像的方法,集成了文本编码器与三个核心组件:用于编码3D CT体数据的新型因果视觉Transformer、用于对齐CT与文本标记的文本-图像Transformer,以及文本条件超分辨率扩散模型。由于3D医学影像领域缺乏直接可比的同类方法,我们采用前沿技术建立基线以验证本方法有效性。GenerateCT在所有关键指标上显著优于这些基线方法。更重要的是,我们通过评估其在多异常分类任务中的效用,探索了GenerateCT的临床应用。首先,基于真实数据集训练多异常分类器建立基线。为进一步评估模型在外部数据集上的泛化能力及零样本场景下对未见提示的性能,我们采用外部数据集训练分类器并设定额外基准。我们开展两项实验:使用GenerateCT为每组数据集合成等量体数据,将训练数据量翻倍。第一项实验表明,联合真实与生成体数据训练分类器时,AP分数提升11%。第二项实验显示,基于未见提示使用真实与生成体数据训练时,AP分数提升7%。此外,GenerateCT支持将合成训练数据集扩展至任意规模。例如,我们生成了100,000个3D CT体数据(为真实数据集量的五倍),并完全基于这些合成体数据训练分类器。令人瞩目的是,该分类器性能较之基于全部可用真实数据训练的分类器高出8%。最后,领域专家评估生成体数据,确认其与文本提示高度对齐。