World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
翻译:世界-动作模型(WAMs)通过将未来视觉预测与动作生成相结合,已成为具身控制领域一种前景广阔的范式。然而,现有大多数WAM依赖逼真的未来预测,这导致推理延迟较高,难以实现机器人的实时部署。这促使我们设计一种更高效的WAM,既能保留未来视觉预测对控制的益处,又能降低其推理成本。我们提出高效WAM,这是一种在保留控制优势的同时降低未来想象成本的世界-动作模型。高效WAM通过从WAN-2.2-5B迁移的紧凑视频专家、稀疏视频潜在令牌以及非对称视频-动作去噪(为视频分配少于动作的采样步数)来提升推理效率。高效WAM并未优化未来分支的视觉保真度,而是将未来视频预测视为动作生成的紧凑引导信号。在RoboTwin 2.0和真实世界操控任务上的综合实验表明,尽管未来预测在视觉上较为粗糙,高效WAM仍能保持强劲的动作性能。在保持竞争性控制能力的同时,我们的10亿参数模型可在物理部署中将每块推理延迟降至约100毫秒,相较于现有WAM实现了30倍的加速。