Access to high-quality and diverse 3D articulated digital human assets is crucial in various applications, ranging from virtual reality to social platforms. Generative approaches, such as 3D generative adversarial networks (GANs), are rapidly replacing laborious manual content creation tools. However, existing 3D GAN frameworks typically rely on scene representations that leverage either template meshes, which are fast but offer limited quality, or volumes, which offer high capacity but are slow to render, thereby limiting the 3D fidelity in GAN settings. In this work, we introduce layered surface volumes (LSVs) as a new 3D object representation for articulated digital humans. LSVs represent a human body using multiple textured mesh layers around a conventional template. These layers are rendered using alpha compositing with fast differentiable rasterization, and they can be interpreted as a volumetric representation that allocates its capacity to a manifold of finite thickness around the template. Unlike conventional single-layer templates that struggle with representing fine off-surface details like hair or accessories, our surface volumes naturally capture such details. LSVs can be articulated, and they exhibit exceptional efficiency in GAN settings, where a 2D generator learns to synthesize the RGBA textures for the individual layers. Trained on unstructured, single-view 2D image datasets, our LSV-GAN generates high-quality and view-consistent 3D articulated digital humans without the need for view-inconsistent 2D upsampling networks.
翻译:高质量且多样化的三维铰接数字人体资产对于从虚拟现实到社交平台等诸多应用至关重要。生成式方法(如三维生成对抗网络GANs)正迅速取代耗时的手动内容创建工具。然而,现有三维GAN框架通常依赖的场景表示要么采用模板网格(速度快但质量有限),要么采用体素(容量高但渲染慢),从而限制了GAN设置中的三维保真度。本文提出分层表面体素(LSVs)作为一种面向铰接数字人体的新型三维物体表示。LSVs通过在传统模板周围使用多层纹理网格层来表征人体。这些层通过alpha合成与快速可微光栅化进行渲染,可被解释为一种将容量分配至模板周围有限厚度流形的体素表示。与难以表现头发或配饰等精细表面细节的传统单层模板不同,我们的表面体素自然捕捉此类细节。LSVs具备铰接能力,并在GAN设置中展现出卓越效率——二维生成器通过学习合成各层的RGBA纹理。基于非结构化单视角二维图像数据集训练,我们的LSV-GAN能够生成高质量且视图一致的三维铰接数字人体,无需使用视图不一致的二维上采样网络。