In this work, we introduce a novel approach for creating controllable dynamics in 3D-generated Gaussians using casually captured reference videos. Our method transfers the motion of objects from reference videos to a variety of generated 3D Gaussians across different categories, ensuring precise and customizable motion transfer. We achieve this by employing blend skinning-based non-parametric shape reconstruction to extract the shape and motion of reference objects. This process involves segmenting the reference objects into motion-related parts based on skinning weights and establishing shape correspondences with generated target shapes. To address shape and temporal inconsistencies prevalent in existing methods, we integrate physical simulation, driving the target shapes with matched motion. This integration is optimized through a displacement loss to ensure reliable and genuine dynamics. Our approach supports diverse reference inputs, including humans, quadrupeds, and articulated objects, and can generate dynamics of arbitrary length, providing enhanced fidelity and applicability. Unlike methods heavily reliant on diffusion video generation models, our technique offers specific and high-quality motion transfer, maintaining both shape integrity and temporal consistency.
翻译:本文提出了一种新方法,通过使用随意拍摄的参考视频为3D生成高斯模型创建可控动力学。该方法能将参考视频中物体的运动迁移至各类生成的3D高斯模型,确保精确且可定制的运动转移。我们通过基于混合蒙皮的非参数形状重建提取参考物体的形状与运动,这一过程涉及根据蒙皮权重将参考物体分割为运动相关部件,并建立与生成目标形状的对应关系。为解决现有方法中普遍存在的形状与时间不一致问题,我们融合物理模拟,利用匹配的运动驱动目标形状,并通过位移损失优化该融合过程,以生成可靠且真实的动力学。我们的方法支持包括人类、四足动物及铰接物体在内的多种参考输入,可生成任意长度的动力学序列,显著提升了保真度与适用性。与严重依赖扩散视频生成模型的方法不同,本技术实现了特定且高质量的运动迁移,同时保持了形状完整性与时间一致性。