Creating high-quality materials in computer graphics is a challenging and time-consuming task, which requires great expertise. To simply this process, we introduce MatFuse, a unified approach that harnesses the generative power of diffusion models to simplify the creation of SVBRDF maps. Our pipeline integrates multiple sources of conditioning, including color palettes, sketches, text, and pictures, for a fine-grained control and flexibility in material synthesis. This design enables the combination of diverse information sources (e.g., sketch + text), enhancing creative possibilities in line with the principle of compositionality. Additionally, we propose a multi-encoder compression model with a two-fold purpose: it improves reconstruction performance by learning a separate latent representation for each map and enables a map-level material editing capabilities. We demonstrate the effectiveness of MatFuse under multiple conditioning settings and explore the potential of material editing. We also quantitatively assess the quality of the generated materials in terms of CLIP-IQA and FID scores. \\ Source code for training MatFuse will be made publically available at https://gvecchio.com/matfuse.
翻译:在计算机图形学中创建高质量材质是一项耗时且富有挑战性的任务,需要具备丰富的专业知识。为简化这一流程,我们提出MatFuse——一种统一方法,利用扩散模型的生成能力来简化SVBRDF贴图的创建过程。我们的管道集成了多重条件控制源,包括调色板、草图、文本和图片,以实现材质合成中的细粒度控制与灵活性。该设计支持多种信息源(如草图+文本)的组合增强,遵循组合性原则提升创作可能性。此外,我们提出一种具有双重目标的多编码器压缩模型:通过为每个贴图学习独立的潜在表征来提升重建性能,同时实现贴图级别的材质编辑能力。我们在多重条件设置下验证了MatFuse的有效性,并探索了材质编辑的潜力。我们还通过CLIP-IQA和FID评分定量评估生成材质的质量。MatFuse的训练源代码将在https://gvecchio.com/matfuse 公开提供。