We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output, without limiting the analysis to narrow sub-areas. It also results in the identification of common templates in the field, which are then compared and contrasted both within pools of similar methods and across lines of research. We provide a breakdown of text-to-image generation into various flavors of image-from-text methods, video-from-text methods, image editing, self-supervised and graph-based approaches. In this discussion, we focus on research papers published at 8 leading machine learning conferences in the years 2016-2022, also incorporating a number of relevant papers not matching the outlined search criteria. The conducted review suggests a significant increase in the number of papers published in the area and highlights research gaps and potential lines of investigation. To our knowledge, this is the first review to systematically look at text-to-image generation from the perspective of "cross-modal generation."
翻译:本文从“跨模态生成”的角度综述了从文本生成视觉数据的研究。这一视角使我们能够将各种处理输入文本并生成视觉输出的方法进行类比,而不将分析局限于狭窄的子领域。同时,这也促进了领域内通用模板的识别,我们随后在同类方法池内部及跨研究路线之间对这些模板进行了比较与对比。我们将文本到图像的生成细分为多种类别:基于文本的图像生成方法、基于文本的视频生成方法、图像编辑、自监督方法以及基于图的方法。在讨论中,我们重点关注2016-2022年间在8个顶级机器学习会议上发表的论文,同时也纳入了部分不符合既定搜索标准的相关论文。本综述表明该领域论文数量显著增长,并指出了研究空白与潜在的研究方向。据我们所知,这是首篇从“跨模态生成”视角系统审视文本到图像生成的综述。