Rigging and skinning clothed human avatars is a challenging task and traditionally requires a lot of manual work and expertise. Recent methods addressing it either generalize across different characters or focus on capturing the dynamics of a single character observed under different pose configurations. However, the former methods typically predict solely static skinning weights, which perform poorly for highly articulated poses, and the latter ones either require dense 3D character scans in different poses or cannot generate an explicit mesh with vertex correspondence over time. To address these challenges, we propose a fully automated approach for creating a fully rigged character with pose-dependent skinning weights, which can be solely learned from multi-view video. Therefore, we first acquire a rigged template, which is then statically skinned. Next, a coordinate-based MLP learns a skinning weights field parameterized over the position in a canonical pose space and the respective pose. Moreover, we introduce our pose- and view-dependent appearance field allowing us to differentiably render and supervise the posed mesh using multi-view imagery. We show that our approach outperforms state-of-the-art while not relying on dense 4D scans.
翻译:为身着衣物的类人化身进行骨骼绑定与蒙皮是一项极具挑战性的任务,传统上需要大量人工操作与专业知识。现有处理方法或跨角色泛化,或专注于捕捉特定角色在不同姿态配置下的动态表现。然而,前者通常仅预测静态蒙皮权重,在高自由度关节姿态下表现欠佳;后者要么需要密集的三维角色姿态扫描数据,要么无法生成具有顶点时序对应关系的显式网格。为解决这些难题,我们提出了一种全自动方法,可创建具有姿态相关蒙皮权重的完整骨架化角色,该方法仅需从多视角视频中学习。为此,我们首先获取一个已绑定骨骼的模板,并为其施加静态蒙皮。随后,基于坐标的多层感知机(MLP)学习一个蒙皮权重场,该场以规范姿态空间中的位置及对应姿态为参数。此外,我们引入姿态与视角相关的外观场,从而能够通过多视角图像对姿态化网格进行可微分渲染与监督。实验表明,本方法在无需密集四维扫描的条件下,性能优于现有最先进技术。