We present Sketch2Colab, which turns storyboard-style 2D sketches into coherent, object-aware 3D multi-human motion with fine-grained control over agents, joints, timing, and contacts. Diffusion-based motion generators offer strong realism but often rely on costly guidance for multi-entity control and degrade under strong conditioning. Sketch2Colab instead learns a sketch-conditioned diffusion prior and distills it into a rectified-flow student in latent space for fast, stable sampling. To make motion follow storyboards closely, we guide the student with differentiable objectives that enforce keyframes, paths, contacts, and physical consistency. Collaborative motion naturally involves discrete changes in interaction, such as converging, forming contact, cooperative transport, or disengaging, and a continuous flow alone struggles to sequence these shifts cleanly. We address this with a lightweight continuous-time Markov chain (CTMC) planner that tracks the active interaction regime and modulates the flow to produce clearer, synchronized coordination in human-object-human motion. Experiments on CORE4D and InterHuman show that Sketch2Colab outperforms baselines in constraint adherence and perceptual quality while sampling substantially faster than diffusion-only alternatives.
翻译:我们提出Sketch2Colab,可将故事板风格的二维草图转化为连贯且具物体感知的三维多人体运动,实现对智能体、关节、时序和接触的精细控制。基于扩散的运动生成器虽具有强真实感,但在多实体控制中常依赖昂贵引导,且在强条件约束下性能退化。Sketch2Colab则学习草图条件扩散先验,并将其蒸馏至潜空间中的修正流学生网络,以实现快速稳定的采样。为使运动紧密贴合故事板,我们通过可微分目标函数引导学生网络,强制执行关键帧、路径、接触及物理一致性。协同运动自然包含相互作用中的离散变化,如汇聚、建立接触、协作搬运或脱离,而单一连续流难以清晰排列这些转换。为此,我们引入轻量级连续时间马尔可夫链(CTMC)规划器,追踪活跃交互模式并调制流生成,从而在人-物体-人运动中产生更清晰、同步的协调动作。在CORE4D和InterHuman上的实验表明,Sketch2Colab在约束遵循度和感知质量上超越基线方法,同时采样速度显著快于纯扩散替代方案。