We study the problem of 3D-aware full-body human generation, aiming at creating animatable human avatars with high-quality textures and geometries. Generally, two challenges remain in this field: i) existing methods struggle to generate geometries with rich realistic details such as the wrinkles of garments; ii) they typically utilize volumetric radiance fields and neural renderers in the synthesis process, making high-resolution rendering non-trivial. To overcome these problems, we propose GETAvatar, a Generative model that directly generates Explicit Textured 3D meshes for animatable human Avatar, with photo-realistic appearance and fine geometric details. Specifically, we first design an articulated 3D human representation with explicit surface modeling, and enrich the generated humans with realistic surface details by learning from the 2D normal maps of 3D scan data. Second, with the explicit mesh representation, we can use a rasterization-based renderer to perform surface rendering, allowing us to achieve high-resolution image generation efficiently. Extensive experiments demonstrate that GETAvatar achieves state-of-the-art performance on 3D-aware human generation both in appearance and geometry quality. Notably, GETAvatar can generate images at 512x512 resolution with 17FPS and 1024x1024 resolution with 14FPS, improving upon previous methods by 2x. Our code and models will be available.
翻译:摘要: 本文研究三维感知全身人体生成问题,旨在创建具有高质量纹理和几何结构的可动画化人体虚拟形象。总体而言,该领域仍面临两大挑战:i) 现有方法难以生成具有丰富逼真细节(如衣物褶皱)的几何结构;ii) 它们在合成过程中通常采用体积辐射场和神经渲染器,使得高分辨率渲染变得困难。为解决上述问题,我们提出GETAvatar——一种直接生成具有显式纹理的三维可动画人体虚拟形象网格的生成式模型,其具备照片级真实外观和精细几何细节。具体而言,我们首先设计了一种具有显式表面建模的铰接式三维人体表示,并通过学习三维扫描数据的二维法线图来增强生成人体的真实表面细节。其次,借助显式网格表示,我们可利用基于光栅化的渲染器执行表面渲染,从而实现高效的高分辨率图像生成。大量实验表明,GETAvatar在外观和几何质量方面的三维感知人体生成任务上均达到了最优性能。值得注意的是,GETAvatar能以17FPS生成512×512分辨率图像、以14FPS生成1024×1024分辨率图像,较先前方法性能提升2倍。代码与模型将公开提供。