We present HAAR, a new strand-based generative model for 3D human hairstyles. Specifically, based on textual inputs, HAAR produces 3D hairstyles that could be used as production-level assets in modern computer graphics engines. Current AI-based generative models take advantage of powerful 2D priors to reconstruct 3D content in the form of point clouds, meshes, or volumetric functions. However, by using the 2D priors, they are intrinsically limited to only recovering the visual parts. Highly occluded hair structures can not be reconstructed with those methods, and they only model the ''outer shell'', which is not ready to be used in physics-based rendering or simulation pipelines. In contrast, we propose a first text-guided generative method that uses 3D hair strands as an underlying representation. Leveraging 2D visual question-answering (VQA) systems, we automatically annotate synthetic hair models that are generated from a small set of artist-created hairstyles. This allows us to train a latent diffusion model that operates in a common hairstyle UV space. In qualitative and quantitative studies, we demonstrate the capabilities of the proposed model and compare it to existing hairstyle generation approaches.
翻译:我们提出HAAR,一种基于发丝的3D人体发型生成模型。具体而言,HAAR根据文本输入生成可直接用于现代计算机图形引擎生产级资产的3D发型。当前基于AI的生成模型利用强大的2D先验知识,以点云、网格或体积函数形式重建3D内容。然而,由于依赖2D先验,这些方法本质上仅能恢复视觉可见部分,高度遮挡的发丝结构无法被重建,且仅能建模"外壳"形态,无法直接用于物理渲染或仿真管线。与此相反,我们首次提出以文本引导的生成方法,采用3D发丝作为底层表示。通过利用2D视觉问答(VQA)系统,我们自动标注由少量艺术家创建发型生成的合成发型模型,从而训练在通用发型UV空间中运行的潜扩散模型。在定性与定量研究中,我们展示了所提模型的能力,并与现有发型生成方法进行了比较。