This paper presents Paint3D, a novel coarse-to-fine generative framework that is capable of producing high-resolution, lighting-less, and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embedded illumination information, which allows the textures to be re-lighted or re-edited within modern graphics pipelines. To achieve this, our method first leverages a pre-trained depth-aware 2D diffusion model to generate view-conditional images and perform multi-view texture fusion, producing an initial coarse texture map. However, as 2D models cannot fully represent 3D shapes and disable lighting effects, the coarse texture map exhibits incomplete areas and illumination artifacts. To resolve this, we train separate UV Inpainting and UVHD diffusion models specialized for the shape-aware refinement of incomplete areas and the removal of illumination artifacts. Through this coarse-to-fine process, Paint3D can produce high-quality 2K UV textures that maintain semantic consistency while being lighting-less, significantly advancing the state-of-the-art in texturing 3D objects.
翻译:本文提出了Paint3D,一种新颖的由粗到精的生成框架,能够基于文本或图像输入,为未纹理化的3D网格生成高分辨率、无光照且多样化的2K UV纹理图。要解决的关键挑战是生成不包含嵌入光照信息的高质量纹理,这使得纹理可以在现代图形管线中进行重新照明或编辑。为此,我们的方法首先利用预训练的深度感知2D扩散模型生成视角条件图像,并执行多视角纹理融合,生成初始的粗略纹理图。然而,由于2D模型无法完全表示3D形状且禁用了光照效果,粗略纹理图存在不完整区域和光照伪影。为解决这一问题,我们训练了专门的UV修复扩散模型和UV高清扩散模型,分别用于形状感知的不完整区域优化和光照伪影去除。通过这个由粗到精的过程,Paint3D能够生成保持语义一致性且无光照的高质量2K UV纹理,显著推动了3D物体纹理生成领域的最新技术水平。