We present Blocks2World, a novel method for 3D scene rendering and editing that leverages a two-step process: convex decomposition of images and conditioned synthesis. Our technique begins by extracting 3D parallelepipeds from various objects in a given scene using convex decomposition, thus obtaining a primitive representation of the scene. These primitives are then utilized to generate paired data through simple ray-traced depth maps. The next stage involves training a conditioned model that learns to generate images from the 2D-rendered convex primitives. This step establishes a direct mapping between the 3D model and its 2D representation, effectively learning the transition from a 3D model to an image. Once the model is fully trained, it offers remarkable control over the synthesis of novel and edited scenes. This is achieved by manipulating the primitives at test time, including translating or adding them, thereby enabling a highly customizable scene rendering process. Our method provides a fresh perspective on 3D scene rendering and editing, offering control and flexibility. It opens up new avenues for research and applications in the field, including authoring and data augmentation.
翻译:我们提出Blocks2World,一种新颖的三维场景渲染与编辑方法,该方法采用两步流程:图像的凸分解与条件合成。该技术首先利用凸分解从给定场景的各类物体中提取三维平行六面体,从而获得场景的基元表示。随后通过简单的光线追踪深度图,利用这些基元生成配对数据。下一阶段训练一个条件模型,使其学会从二维渲染的凸基元生成图像。此步骤建立了三维模型与其二维表示之间的直接映射,有效学习了从三维模型到图像的转换过程。模型完全训练后,能够对新颖场景与编辑场景的合成提供卓越控制。通过测试时对基元进行操作(包括平移或添加基元)实现此效果,从而赋予场景渲染过程高度可定制性。我们的方法为三维场景渲染与编辑提供了全新视角,兼具控制力与灵活性,为该领域的创作与数据增强等研究方向与应用开辟了新途径。