Text to image generation methods (T2I) are widely popular in generating art and other creative artifacts. While visual hallucinations can be a positive factor in scenarios where creativity is appreciated, such artifacts are poorly suited for cases where the generated image needs to be grounded in complex natural language without explicit visual elements. In this paper, we propose to strengthen the consistency property of T2I methods in the presence of natural complex language, which often breaks the limits of T2I methods by including non-visual information, and textual elements that require knowledge for accurate generation. To address these phenomena, we propose a Natural Language to Verified Image generation approach (NL2VI) that converts a natural prompt into a visual prompt, which is more suitable for image generation. A T2I model then generates an image for the visual prompt, which is then verified with VQA algorithms. Experimentally, aligning natural prompts with image generation can improve the consistency of the generated images by up to 11% over the state of the art. Moreover, improvements can generalize to challenging domains like cooking and DIY tasks, where the correctness of the generated image is crucial to illustrate actions.
翻译:文本到图像生成方法(T2I)在艺术创作及其他创意性作品生成中广受欢迎。虽然视觉幻觉在重视创意的场景中可能成为积极因素,但当生成图像需要基于复杂自然语言且缺乏显式视觉元素时,这类产物往往难以胜任。本文旨在提升T2I方法在处理复杂自然语言时的一致性——这类语言常因包含非视觉信息及依赖知识才能准确生成文本元素,而突破T2I方法的限制。为解决上述问题,我们提出自然语言到验证图像生成方法(NL2VI),该方法将自然提示转换为更适于图像生成的视觉提示。随后,T2I模型为视觉提示生成图像,并通过VQA算法进行验证。实验表明,将自然提示与图像生成对齐后,生成图像的一致性较现有最优方法提升高达11%。此外,该改进可泛化至烹饪、DIY任务等挑战性领域——在这些场景中,生成图像的准确性对动作演示至关重要。