Generating visual layouts is an essential ingredient of graphic design. The ability to condition layout generation on a partial subset of component attributes is critical to real-world applications that involve user interaction. Recently, diffusion models have demonstrated high-quality generative performances in various domains. However, it is unclear how to apply diffusion models to the natural representation of layouts which consists of a mix of discrete (class) and continuous (location, size) attributes. To address the conditioning layout generation problem, we introduce DLT, a joint discrete-continuous diffusion model. DLT is a transformer-based model which has a flexible conditioning mechanism that allows for conditioning on any given subset of all the layout component classes, locations, and sizes. Our method outperforms state-of-the-art generative models on various layout generation datasets with respect to different metrics and conditioning settings. Additionally, we validate the effectiveness of our proposed conditioning mechanism and the joint continuous-diffusion process. This joint process can be incorporated into a wide range of mixed discrete-continuous generative tasks.
翻译:摘要:视觉布局生成是图形设计中的关键要素。在真实世界的人机交互应用中,根据部分组件属性的子集对布局生成进行条件化处理的能力至关重要。近年来,扩散模型已在多个领域展现出高质量的生成性能。然而,如何将扩散模型应用于由离散(类别)和连续(位置、尺寸)属性混合构成的布局自然表示,仍不明确。为解决条件化布局生成问题,我们提出DLT——一种联合离散-连续扩散模型。DLT基于Transformer架构,具备灵活的条件化机制,可对布局组件类别、位置和尺寸的任意子集进行条件控制。在多个布局生成数据集上,我们的方法在不同评价指标和条件设置下均优于现有最优生成模型。此外,我们验证了所提条件化机制与联合连续-离散扩散过程的有效性。该联合过程可推广至多种混合离散-连续生成任务。