Animation of 2D hand-drawn sketches provides an effective medium for visual communication. However, these sketches pose challenges, particularly in handling occlusions and accurately mapping motion. While 3D animation naturally addresses these challenges, estimating 3D motion remains a very complex task. Recent approaches to converting 2D sketches to 3D animations have mainly focused on specific types of motion, such as bipedal movements and facial expressions. We propose Sketch2Motion, a diffusion-guided framework for skeleton-based motion synthesis that combines classical character animation pipelines with deep generative priors. Our method represents motion using skeletal transformations, which are propagated to mesh deformations via linear blend skinning. To guide the resulting animation toward realistic and semantically meaningful motion, we integrate a text-to-video diffusion model via motion-aware score-distillation sampling (MoSDS), enabling optimization without paired motion data. Additionally, we apply physics-inspired smoothness, topological, and contact constraints to stabilize optimization and preserve motion plausibility. Further, we integrate a spring-mass simulator to introduce secondary motion effects. The proposed framework is generalized, fully differentiable, modular, and compatible with biped, quadruped, and non-living articulated characters. Experiments demonstrate that our approach produces temporally coherent, text-aligned animations that outperform baseline motion transfer methods that lack generative priors or explicit physical constraints. We will make our code and dataset publicly available.
翻译:二维手绘草图的动画生成为视觉交流提供了有效媒介。然而,此类草图面临挑战,尤其在处理遮挡和精确映射运动方面。尽管三维动画天然解决了这些问题,但三维运动估计仍是一项极为复杂的任务。近期将二维草图转换为三维动画的方法主要聚焦于特定运动类型,如双足运动与面部表情。我们提出Sketch2Motion——一种基于扩散引导的骨骼运动合成框架,该框架将经典角色动画管线与深度生成先验相结合。该方法利用骨骼变换表示运动,并通过线性混合蒙皮将运动映射至网格变形。为引导生成动画朝向符合物理规律与语义合理性的运动,我们通过运动感知分数蒸馏采样(MoSDS)集成文本到视频扩散模型,从而无需配对运动数据即可实现优化。此外,我们应用物理启发的平滑性约束、拓扑约束与接触约束以稳定优化并保持运动合理性。进一步地,我们集成弹簧-质点模拟器以引入次级运动效果。所提框架具有通用性、全可微性、模块化特性,并兼容双足、四足及非生命体铰接角色。实验表明,我们的方法能生成时间连贯且与文本对齐的动画,其性能优于缺乏生成先验或显式物理约束的基线运动迁移方法。我们将公开代码与数据集。