The generative modeling landscape has experienced tremendous growth in recent years, particularly in generating natural images and art. Recent techniques have shown impressive potential in creating complex visual compositions while delivering impressive realism and quality. However, state-of-the-art methods have been focusing on the narrow domain of natural images, while other distributions remain unexplored. In this paper, we introduce the problem of text-to-figure generation, that is creating scientific figures of papers from text descriptions. We present FigGen, a diffusion-based approach for text-to-figure as well as the main challenges of the proposed task. Code and models are available at https://github.com/joanrod/figure-diffusion
翻译:生成式建模领域近年来经历了巨大增长,特别是在自然图像和艺术生成方面。最新技术展示了在创建复杂视觉构图方面的惊人潜力,同时实现了令人印象深刻的真实感和质量。然而,最先进的方法一直聚焦于自然图像这一狭窄领域,而其他分布仍未得到探索。在本文中,我们引入了文本到图形生成问题,即根据文本描述生成论文的科学图形。我们提出FigGen,一种基于扩散的文本到图形方法,并阐述了该任务的主要挑战。代码和模型可在https://github.com/joanrod/figure-diffusion获取。