Much progress has been made in reconstructing garments from an image or a video. However, none of existing works meet the expectations of digitizing high-quality animatable dynamic garments that can be adjusted to various unseen poses. In this paper, we propose the first method to recover high-quality animatable dynamic garments from monocular videos without depending on scanned data. To generate reasonable deformations for various unseen poses, we propose a learnable garment deformation network that formulates the garment reconstruction task as a pose-driven deformation problem. To alleviate the ambiguity estimating 3D garments from monocular videos, we design a multi-hypothesis deformation module that learns spatial representations of multiple plausible deformations. Experimental results on several public datasets demonstrate that our method can reconstruct high-quality dynamic garments with coherent surface details, which can be easily animated under unseen poses. The code will be provided for research purposes.
翻译:从单张图像或视频中重建服装已取得显著进展。然而,现有方法均无法满足对可调整至各种未见姿态的高质量可动画动态服装数字化的期望。本文提出首种无需依赖扫描数据、从单目视频恢复高质量可动画动态服装的方法。为生成适应各类未见姿态的合理形变,我们提出了可学习的服装形变网络,将服装重建任务建模为姿态驱动的形变问题。为缓解单目视频中三维服装估计的歧义性,我们设计了多假设形变模块,学习多种合理形变的空间表征。多个公开数据集上的实验结果表明,本方法可重建具有连贯表面细节的高质量动态服装,且能轻松在未见姿态下实现动画化。相关代码将面向研究目的开放。