Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational resources and training time, and require additional human annotations such as predefined parametric models, camera poses, and key points, limiting their generalizability. We propose Synergistic Shape and Skeleton Optimization (S3O), a novel two-phase method that forgoes these prerequisites and efficiently learns parametric models including visible shapes and underlying skeletons. Conventional strategies typically learn all parameters simultaneously, leading to interdependencies where a single incorrect prediction can result in significant errors. In contrast, S3O adopts a phased approach: it first focuses on learning coarse parametric models, then progresses to motion learning and detail addition. This method substantially lowers computational complexity and enhances robustness in reconstruction from limited viewpoints, all without requiring additional annotations. To address the current inadequacies in 3D reconstruction from monocular video benchmarks, we collected the PlanetZoo dataset. Our experimental evaluations on standard benchmarks and the PlanetZoo dataset affirm that S3O provides more accurate 3D reconstruction, and plausible skeletons, and reduces the training time by approximately 60% compared to the state-of-the-art, thus advancing the state of the art in dynamic object reconstruction.
翻译:从单一单目视频中重建动态铰接物体极具挑战性,需要从有限视角中联合估计形状、运动与相机参数。现有方法通常要求大量计算资源与训练时间,并依赖预定义参数化模型、相机姿态与关键点等额外人工标注,限制了其泛化能力。我们提出协同形状与骨架优化(S3O),一种不依赖上述先验条件的新型双阶段方法,可高效学习可见形状与潜在骨架等参数化模型。传统策略通常同步学习所有参数,导致单次错误预测可能引发显著误差的相互依赖问题。与之相反,S3O采用分阶段方法:首先专注于粗粒度参数化模型的学习,随后推进至运动学习与细节补充。该方法大幅降低计算复杂度,增强有限视角下重建的鲁棒性,且无需额外标注。针对当前单目视频三维重建基准的不足,我们收集了PlanetZoo数据集。在标准基准及PlanetZoo数据集上的实验评估证实,S3O提供了更精确的三维重建与合理的骨架结构,并将训练时间相比现有最优方法降低约60%,从而推动了动态物体重建领域的技术发展。