How humans can efficiently and effectively acquire images has always been a perennial question. A typical solution is text-to-image retrieval from an existing database given the text query; however, the limited database typically lacks creativity. By contrast, recent breakthroughs in text-to-image generation have made it possible to produce fancy and diverse visual content, but it faces challenges in synthesizing knowledge-intensive images. In this work, we rethink the relationship between text-to-image generation and retrieval and propose a unified framework in the context of Multimodal Large Language Models (MLLMs). Specifically, we first explore the intrinsic discriminative abilities of MLLMs and introduce a generative retrieval method to perform retrieval in a training-free manner. Subsequently, we unify generation and retrieval in an autoregressive generation way and propose an autonomous decision module to choose the best-matched one between generated and retrieved images as the response to the text query. Additionally, we construct a benchmark called TIGeR-Bench, including creative and knowledge-intensive domains, to standardize the evaluation of unified text-to-image generation and retrieval. Extensive experimental results on TIGeR-Bench and two retrieval benchmarks, i.e., Flickr30K and MS-COCO, demonstrate the superiority and effectiveness of our proposed method.
翻译:人类如何高效且有效地获取图像始终是一个长期性问题。典型的解决方案是基于文本查询从现有数据库中检索图像,但有限的数据库往往缺乏创造性。相比之下,文本到图像生成领域的最新突破使得生成新颖多样的视觉内容成为可能,但在合成知识密集型图像时仍面临挑战。在本工作中,我们重新审视了文本到图像生成与检索之间的关系,并在多模态大语言模型(MLLMs)框架下提出了一个统一框架。具体而言,我们首先探索了MLLMs的内在判别能力,并引入了一种无需训练的生成式检索方法。随后,我们以自回归生成方式统一了生成与检索,并提出了一种自主决策模块,用于在生成图像与检索图像之间选择最佳匹配结果作为对文本查询的响应。此外,我们构建了名为TIGeR-Bench的基准测试集,涵盖创意型与知识密集型领域,以标准化统一文本到图像生成与检索的评估。在TIGeR-Bench以及两个检索基准测试集(Flickr30K和MS-COCO)上的广泛实验结果证明了所提出方法的优越性与有效性。