Spatial control is a core capability in controllable image generation. Advancements in layout-guided image generation have shown promising results on in-distribution (ID) datasets with similar spatial configurations. However, it is unclear how these models perform when facing out-of-distribution (OOD) samples with arbitrary, unseen layouts. In this paper, we propose LayoutBench, a diagnostic benchmark for layout-guided image generation that examines four categories of spatial control skills: number, position, size, and shape. We benchmark two recent representative layout-guided image generation methods and observe that the good ID layout control may not generalize well to arbitrary layouts in the wild (e.g., objects at the boundary). Next, we propose IterInpaint, a new baseline that generates foreground and background regions in a step-by-step manner via inpainting, demonstrating stronger generalizability than existing models on OOD layouts in LayoutBench. We perform quantitative and qualitative evaluation and fine-grained analysis on the four LayoutBench skills to pinpoint the weaknesses of existing models. Lastly, we show comprehensive ablation studies on IterInpaint, including training task ratio, crop&paste vs. repaint, and generation order. Project website: https://layoutbench.github.io
翻译:论文摘要:空间控制是可控图像生成的核心能力。布局引导图像生成的最新进展已在具有相似空间配置的分布内(ID)数据集上展现出可喜成果。然而,当面对包含任意未见布局的分布外(OOD)样本时,这些模型的表现尚不明确。本文提出了LayoutBench——针对布局引导图像生成的诊断基准,该基准考察了四种空间控制技能:数量、位置、尺寸和形状。我们对两类近期代表性布局引导图像生成方法进行基准测试后发现,优异的ID布局控制能力可能无法有效泛化至真实场景中的任意布局(例如边界处的物体)。为此,我们提出了IterInpaint,一种通过渐进式修补生成前景与背景区域的新基线方法,其在LayoutBench的OOD布局上展现出比现有模型更强的泛化能力。我们通过定性与定量评估,并结合针对四种LayoutBench技能的细粒度分析,精准定位了现有模型的缺陷。最后,我们对IterInpaint进行了全面的消融研究,涵盖训练任务比例、裁剪粘贴vs重绘策略以及生成顺序等关键因素。项目网站:https://layoutbench.github.io