Despite transformative advances in generative motion synthesis, real-time interactive motion control remains dominated by traditional techniques. In this work, we identify two key challenges in bridging research and production: 1) Real-time scalability: Industry applications demand real-time generation of a vast repertoire of motion skills, while generative methods exhibit significant degradation in quality and scalability under real-time computation constraints, and 2) Integration: Industry applications demand fine-grained multi-modal control involving velocity commands, style selection, and precise keyframes, a need largely unmet by existing text- or tag-driven models. To overcome these limitations, we introduce MotionBricks: a large-scale, real-time generative framework with a two-fold solution. First, we propose a large-scale modular latent generative backbone tailored for robust real-time motion generation, effectively modeling a dataset of over 350,000 motion clips with a single model. Second, we introduce smart primitives that provide a unified, robust, and intuitive interface for authoring both navigation and object interaction. Applications can be designed in a plug-and-play manner like assembling bricks without expert animation knowledge. Quantitatively, we show that MotionBricks produces state-of-the-art motion quality on open-source and proprietary datasets of various scales, while also achieving a real-time throughput of 15,000 FPS with 2ms latency. We demonstrate the flexibility and robustness of MotionBricks in a complete production-level animation demo, covering navigation and object-scene interaction across various styles with a unified model. To showcase our framework's application beyond animation, we deploy MotionBricks on the Unitree G1 humanoid robot to demonstrate its flexibility and generalization for real-time robotic control.
翻译:尽管生成式动作合成取得了变革性进展,实时交互式动作控制仍由传统技术主导。本文揭示了连接研究与实际应用的两大核心挑战:1)实时可扩展性:工业应用要求实时生成海量动作技能库,而生成式方法在实时计算约束下存在质量和可扩展性显著下降的问题;2)集成性:工业应用需要涉及速度指令、风格选择和精确关键帧的细粒度多模态控制,这一需求尚未被现有文本或标签驱动模型充分满足。为突破这些局限,我们提出MotionBricks——一个兼具双重解决方案的大规模实时生成框架。首先,我们构建了专为鲁棒实时动作生成设计的大规模模块化潜变量生成主干,通过单一模型有效建模了包含超过35万个动作片段的数据集。其次,我们引入智能原语,为导航和物体交互的创作提供了统一、鲁棒且直观的接口。应用可实现类积木拼接式的即插即用设计,无需动画专业知识。定量实验表明,MotionBricks在多种规模的开源和专有数据集上均取得了当前最优的动作质量,同时实现15000 FPS的实时吞吐量与2ms延迟。我们通过完整的生产级动画演示验证了MotionBricks的灵活性和鲁棒性,展示了单一模型对各种风格下导航与物体场景交互的统一支持。为展现框架超越动画领域的应用价值,我们在宇树G1人形机器人上部署MotionBricks,验证了其在实时机器人控制中的灵活性与泛化能力。