Recent advances in implicit function-based approaches have shown promising results in 3D human reconstruction from a single RGB image. However, these methods are not sufficient to extend to more general cases, often generating dragged or disconnected body parts, particularly for animated characters. We argue that these limitations stem from the use of the existing point-level 3D shape representation, which lacks holistic 3D context understanding. Voxel-based reconstruction methods are more suitable for capturing the entire 3D space at once, however, these methods are not practical for high-resolution reconstructions due to their excessive memory usage. To address these challenges, we introduce Tri-directional Implicit Function (TIFu), which is a vector-level representation that increases global 3D consistencies while significantly reducing memory usage compared to voxel representations. We also introduce a new algorithm in 3D reconstruction at an arbitrary resolution by aggregating vectors along three orthogonal axes, resolving inherent problems with regressing fixed dimension of vectors. Our approach achieves state-of-the-art performances in both our self-curated character dataset and the benchmark 3D human dataset. We provide both quantitative and qualitative analyses to support our findings.
翻译:基于隐式函数的方法近期在单张RGB图像的三维人体重建中展现出令人瞩目的进展。然而,这些方法难以推广至更具通用性的场景,特别是在动画角色重建中常产生拖拽或断裂的人体部位。我们认为上述局限性源于现有基于点级的三维形状表征方式缺乏对整体三维场景的全局理解。基于体素的重建方法虽能同时捕获整个三维空间,但受限于过高的内存消耗而无法实现高分辨率重建。为解决上述问题,我们提出三向隐式函数(TIFu)——一种向量级表征方法,在显著降低内存开销的同时增强全局三维一致性。通过沿三个正交轴向聚合向量,我们进一步提出任意分辨率下的三维重建新算法,解决了向量固定维度回归的固有问题。本方法在自建角色数据集与标准三维人体数据集中均达到最优性能,并通过定量与定性分析验证了研究成果。