Creating expressive, diverse and high-quality 3D avatars from highly customized text descriptions and pose guidance is a challenging task, due to the intricacy of modeling and texturing in 3D that ensure details and various styles (realistic, fictional, etc). We present AvatarVerse, a stable pipeline for generating expressive high-quality 3D avatars from nothing but text descriptions and pose guidance. In specific, we introduce a 2D diffusion model conditioned on DensePose signal to establish 3D pose control of avatars through 2D images, which enhances view consistency from partially observed scenarios. It addresses the infamous Janus Problem and significantly stablizes the generation process. Moreover, we propose a progressive high-resolution 3D synthesis strategy, which obtains substantial improvement over the quality of the created 3D avatars. To this end, the proposed AvatarVerse pipeline achieves zero-shot 3D modeling of 3D avatars that are not only more expressive, but also in higher quality and fidelity than previous works. Rigorous qualitative evaluations and user studies showcase AvatarVerse's superiority in synthesizing high-fidelity 3D avatars, leading to a new standard in high-quality and stable 3D avatar creation. Our project page is: https://avatarverse3d.github.io
翻译:从高度定制化的文本描述和姿态引导中创建富有表现力、多样化且高质量的三维虚拟人是一项具有挑战性的任务,这源于三维建模与纹理贴图在确保细节及多种风格(写实、虚构等)上的复杂性。我们提出AvatarVerse,一个仅凭文本描述和姿态引导即可生成富有表现力的高质量三维虚拟人的稳定流水线。具体而言,我们引入一个基于DensePose信号条件化的二维扩散模型,通过二维图像建立虚拟人的三维姿态控制,从而增强部分观测场景下的视角一致性。该模型解决了著名的雅努斯问题,并显著稳定了生成过程。此外,我们提出渐进式高分辨率三维合成策略,大幅提升了所创建三维虚拟人的质量。最终,所提出的AvatarVerse流水线实现了三维虚拟人的零样本三维建模,其生成结果不仅比先前工作更具表现力,而且在质量和保真度上更胜一筹。严格的定性评估和用户研究展示了AvatarVerse在合成高保真三维虚拟人方面的优越性,为高质量稳定的三维虚拟人创建树立了新标准。我们的项目页面为:https://avatarverse3d.github.io