We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured in-the-wild image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diffusion models as priors for optimization. First, we extend a foundational model with a lightweight expression-aware and ID-aware architecture, and create 2D priors for geometry and texture generation, via fine-tuning only 0.2% of its available training parameters. Then, we jointly leverage a neural parametric representation for the expressions of each subject and a multi-stage generation of highly detailed geometry and albedo texture. This combination of strong face identity embeddings and our neural representation enables accurate reconstruction of not only facial features but also accessories and hair and can be meshed to provide render-ready assets for gaming and telepresence. Our results achieve an unprecedented level of identity-consistent and high-quality texture and geometry generation, generalizing to a ``world'' of unseen 3D identities, without relying on large 3D captured datasets of human assets.
翻译:我们提出ID-to-3D方法,该方法能够从单张随意捕捉的真实场景人物图像出发,生成具有解耦表情、且受身份与文本引导的3D人头模型。我们方法的基础立足于组合性原理,并利用面向特定任务的2D扩散模型作为优化的先验。首先,我们通过仅微调其可用训练参数的0.2%,扩展了一个基础模型,采用轻量级的表情感知与身份感知架构,并创建了用于几何与纹理生成的2D先验。随后,我们联合利用针对每个主体的表情神经参数化表示,以及高度细节化的几何与反照率纹理的多阶段生成方法。这种强人脸身份嵌入与我们的神经表示的结合,不仅能够精确重建面部特征,还能重建配饰与头发,并可被网格化以提供适用于游戏与远程呈现的渲染就绪资产。我们的结果实现了前所未有的身份一致性及高质量的纹理与几何生成水平,能够泛化到一个由未见过的3D身份构成的“世界”,而无需依赖大规模捕获的人类资产3D数据集。