With the advent of depth-to-image diffusion models, text-guided generation, editing, and transfer of realistic textures are no longer difficult. However, due to the limitations of pre-trained diffusion models, they can only create low-resolution, inconsistent textures. To address this issue, we present the High-definition Consistency Texture Model (HCTM), a novel method that can generate high-definition and consistent textures for 3D meshes according to the text prompts. We achieve this by leveraging a pre-trained depth-to-image diffusion model to generate single viewpoint results based on the text prompt and a depth map. We fine-tune the diffusion model with Parameter-Efficient Fine-Tuning to quickly learn the style of the generated result, and apply the multi-diffusion strategy to produce high-resolution and consistent results from different viewpoints. Furthermore, we propose a strategy that prevents the appearance of noise on the textures caused by backpropagation. Our proposed approach has demonstrated promising results in generating high-definition and consistent textures for 3D meshes, as demonstrated through a series of experiments.
翻译:随着深度到图像扩散模型的出现,文本引导的逼真纹理生成、编辑和迁移已不再困难。然而,由于预训练扩散模型的局限性,它们只能生成低分辨率、不一致的纹理。为解决这一问题,我们提出了高清一致性纹理模型(HCTM),一种能够根据文本提示为三维网格生成高清且一致纹理的新方法。我们通过利用预训练的深度到图像扩散模型,基于文本提示和深度图生成单一视角结果来实现这一点。我们使用参数高效微调对扩散模型进行微调,以快速学习生成结果的风格,并应用多扩散策略从不同视角生成高分辨率且一致的结果。此外,我们提出了一种防止反向传播导致纹理出现噪声的策略。通过一系列实验,我们提出的方法在生成三维网格的高清且一致纹理方面展示了有前景的结果。