We introduce ``Idea to Image,'' a system that enables multimodal iterative self-refinement with GPT-4V(ision) for automatic image design and generation. Humans can quickly identify the characteristics of different text-to-image (T2I) models via iterative explorations. This enables them to efficiently convert their high-level generation ideas into effective T2I prompts that can produce good images. We investigate if systems based on large multimodal models (LMMs) can develop analogous multimodal self-refinement abilities that enable exploring unknown models or environments via self-refining tries. Idea2Img cyclically generates revised T2I prompts to synthesize draft images, and provides directional feedback for prompt revision, both conditioned on its memory of the probed T2I model's characteristics. The iterative self-refinement brings Idea2Img various advantages over vanilla T2I models. Notably, Idea2Img can process input ideas with interleaved image-text sequences, follow ideas with design instructions, and generate images of better semantic and visual qualities. The user preference study validates the efficacy of multimodal iterative self-refinement on automatic image design and generation.
翻译:我们提出“Idea to Image”系统,该系统利用GPT-4V(ision)实现多模态迭代自优化,从而完成自动图像设计与生成。人类能够通过迭代探索快速识别不同文生图(T2I)模型的特征,进而高效地将高层次生成构想转化为能生成优质图像的有效T2I提示词。我们探究基于大型多模态模型(LMMs)的系统是否能够发展出类似的多模态自优化能力,通过自优化尝试来探索未知模型或环境。Idea2Img循环生成修正后的T2I提示词以合成草稿图像,并提供方向性反馈以优化提示词,这一过程均基于其对所探测T2I模型特征的记忆。相较于原始T2I模型,迭代自优化赋予Idea2Img多项优势。值得注意的是,Idea2Img能够处理包含交错图文序列的输入构想、遵循带有设计指令的构想,并生成语义和视觉质量更优的图像。用户偏好研究验证了多模态迭代自优化在自动图像设计与生成中的有效性。