In recent years, significant progress has been made in the development of text- to-image generation models. However, these models still face limitations when it comes to achieving full controllability during the generation process. Often, spe- cific training or the use of limited models is required, and even then, they have certain restrictions. To address these challenges, A two-stage method that effec- tively combines controllability and high quality in the generation of images is proposed. This approach leverages the expertise of pre-trained models to achieve precise control over the generated images, while also harnessing the power of diffusion models to achieve state-of-the-art quality. By separating controllability from high quality, This method achieves outstanding results. It is compatible with both latent and image space diffusion models, ensuring versatility and flexibil- ity. Moreover, This approach consistently produces comparable outcomes to the current state-of-the-art methods in the field. Overall, This proposed method rep- resents a significant advancement in text-to-image generation, enabling improved controllability without compromising on the quality of the generated images.
翻译:近年来,文本到图像生成模型取得了显著进展。然而,这些模型在生成过程中实现完全可控性方面仍面临局限,通常需要特定训练或使用受限模型,即便如此仍存在一定限制。为解决这些问题,本文提出了一种两阶段方法,有效结合了图像生成中的可控性与高质量。该方法利用预训练模型的专长实现对生成图像的精确控制,同时借助扩散模型获得最先进的质量。通过将可控性与高质量分离,本方法取得了卓越成效。它兼容潜空间和图像空间扩散模型,确保了通用性与灵活性。此外,该方法持续产生与当前领域最先进方法相当的结果。总体而言,所提出的方法代表了文本到图像生成的重要进步,在不牺牲生成图像质量的前提下实现了可控性的提升。