It is now possible to reconstruct dynamic human motion and shape from a sparse set of cameras using Neural Radiance Fields (NeRF) driven by an underlying skeleton. However, a challenge remains to model the deformation of cloth and skin in relation to skeleton pose. Unlike existing avatar models that are learned implicitly or rely on a proxy surface, our approach is motivated by the observation that different poses necessitate unique frequency assignments. Neglecting this distinction yields noisy artifacts in smooth areas or blurs fine-grained texture and shape details in sharp regions. We develop a two-branch neural network that is adaptive and explicit in the frequency domain. The first branch is a graph neural network that models correlations among body parts locally, taking skeleton pose as input. The second branch combines these correlation features to a set of global frequencies and then modulates the feature encoding. Our experiments demonstrate that our network outperforms state-of-the-art methods in terms of preserving details and generalization capabilities.
翻译:如今,凭借底层骨架驱动的神经辐射场(NeRF),可以从稀疏相机视角重建动态人体运动与形状。然而,布料与皮肤随骨架姿态的形变建模仍是一大挑战。不同于现有隐式学习或依赖代理表面的虚拟化身模型,我们的方法源自以下观察:不同姿态需要独特的频域分配。忽略这一差异会在平滑区域产生噪声伪影,或在尖锐区域模糊精细纹理与形状细节。我们开发了一种在频域中具有自适应性和显式性的双分支神经网络。第一分支是以骨架姿态为输入的图神经网络,用于局部建模身体部位间的关联。第二分支将这些关联特征整合为全局频率集合,进而调制特征编码。实验表明,我们的网络在细节保持和泛化能力方面均优于现有最先进方法。