Neural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating a 3D human avatar with a specific identity and artistic style that can be easily animated. Our proposed method, AvatarCraft, addresses this challenge by using diffusion models to guide the learning of geometry and texture for a neural avatar based on a single text prompt. We carefully design the optimization framework of neural implicit fields, including a coarse-to-fine multi-bounding box training strategy, shape regularization, and diffusion-based constraints, to produce high-quality geometry and texture. Additionally, we make the human avatar animatable by deforming the neural implicit field with an explicit warping field that maps the target human mesh to a template human mesh, both represented using parametric human models. This simplifies animation and reshaping of the generated avatar by controlling pose and shape parameters. Extensive experiments on various text descriptions show that AvatarCraft is effective and robust in creating human avatars and rendering novel views, poses, and shapes. Our project page is: \url{https://avatar-craft.github.io/}.
翻译:神经隐式场在表示三维场景和生成高质量新视角方面具有强大能力,但利用此类隐式表示创建具有特定身份和艺术风格且易于动画化的人体化身仍具挑战性。我们提出的方法AvatarCraft通过利用扩散模型基于单一文本提示指导神经化身几何与纹理的学习来解决这一挑战。我们精心设计了神经隐式场的优化框架,包括从粗到细的多包围盒训练策略、形状正则化及基于扩散的约束,以生成高质量的几何与纹理。此外,我们通过显式变形场将神经隐式场变形为目标人体网格到模板人体网格的映射(两者均使用参数化人体模型表示),从而使人体化身具备可动画性。这简化了通过控制姿态和形状参数对生成化身进行动画化与重塑的过程。对多种文本描述的广泛实验表明,AvatarCraft在创建人体化身及渲染新视角、姿态和形状方面具有有效性和鲁棒性。项目页面:\url{https://avatar-craft.github.io/}。