We present Kandinsky 3.0, a large-scale text-to-image generation model based on latent diffusion, continuing the series of text-to-image Kandinsky models and reflecting our progress to achieve higher quality and realism of image generation. Compared to previous versions of Kandinsky 2.x, Kandinsky 3.0 leverages a two times larger U-Net backbone, a ten times larger text encoder and removes diffusion mapping. We describe the architecture of the model, the data collection procedure, the training technique, and the production system of user interaction. We focus on the key components that, as we have identified as a result of a large number of experiments, had the most significant impact on improving the quality of our model compared to the others. By our side-by-side comparisons, Kandinsky becomes better in text understanding and works better on specific domains. Project page: https://ai-forever.github.io/Kandinsky-3
翻译:我们提出 Kandinsky 3.0,一个基于潜在扩散的大规模文本到图像生成模型,延续了 Kandinsky 系列文本到图像模型的发展,并反映了我们在实现更高图像生成质量与真实感方面的进展。相较于先前版本的 Kandinsky 2.x,Kandinsky 3.0 利用了两倍大的 U-Net 主干网络、十倍大的文本编码器,并移除了扩散映射。我们描述了模型架构、数据收集流程、训练技术以及用户交互的生产系统。我们重点介绍了基于大量实验所识别出的关键组件,这些组件相较于其他因素,对提升我们模型质量产生了最显著的影响。通过我们的逐对比较,Kandinsky 在文本理解方面表现更优,并在特定领域内效果更好。项目页面:https://ai-forever.github.io/Kandinsky-3