We present HeadEvolver, a novel framework to generate stylized head avatars from text guidance. HeadEvolver uses locally learnable mesh deformation from a template head mesh, producing high-quality digital assets for detail-preserving editing and animation. To tackle the challenges of lacking fine-grained and semantic-aware local shape control in global deformation through Jacobians, we introduce a trainable parameter as a weighting factor for the Jacobian at each triangle to adaptively change local shapes while maintaining global correspondences and facial features. Moreover, to ensure the coherence of the resulting shape and appearance from different viewpoints, we use pretrained image diffusion models for differentiable rendering with regularization terms to refine the deformation under text guidance. Extensive experiments demonstrate that our method can generate diverse head avatars with an articulated mesh that can be edited seamlessly in 3D graphics software, facilitating downstream applications such as more efficient animation with inherited blend shapes and semantic consistency.
翻译:摘要:我们提出HeadEvolver,一种新颖的框架,通过文本引导生成风格化头部头像。HeadEvolver利用模板头部网格的局部可学习网格变形,生成高质量的数字化资产,支持细节保持的编辑与动画。为解决全局变形中通过雅可比矩阵缺乏细粒度、语义感知的局部形状控制难题,我们引入可训练参数作为每个三角形的雅可比权重因子,在保持全局对应关系与面部特征的同时自适应调整局部形状。此外,为确保不同视角下结果形状与外观的一致性,我们利用预训练图像扩散模型进行可微渲染,并引入正则化项以在文本引导下优化变形。大量实验表明,我们的方法可生成多样化头部头像及具有关节结构的网格,该网格可在3D图形软件中无缝编辑,从而驱动下游应用(如通过继承混合形状与语义一致性实现更高效的动画制作)。