Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only video rollouts are not directly actionable: they leave metric 3D motion, contact geometry, and fine-grained spatial constraints under-specified, making action grounding ambiguous. Meanwhile, scaling action supervision across diverse tasks and embodiments remains costly. We present PointAction, a framework that bridges video predictions to robot actions through explicit point-based 4D modeling. PointAction fine-tunes a foundation video generation model to jointly predict future RGB frames and dynamic 3D pointmaps, producing temporally consistent 3D motion of task-relevant scene geometry. These point dynamics serve as a structured, embodiment-agnostic action interface, which a diffusion-based action decoder maps to executable robot actions. By using metric 3D point dynamics as the interface between video prediction and control, PointAction reduces the ambiguity of RGB-only action grounding and supports transfer across tasks and embodiments with limited action supervision. Experiments show that PointAction achieves state-of-the-art 4D generation quality on robot scenes, outperforms existing baselines in simulation, and generalizes to two real robot arms unseen during pretraining.
翻译:视频-行为模型利用预训练视频扩散模型捕捉的广泛视觉动态,为通用型机器人操作提供了可行路径。然而,仅基于RGB的视频推理并不具备直接操作性:这类方法对度量三维运动、接触几何及细微空间约束的刻画不足,导致行为锚定存在模糊性。与此同时,跨不同任务与实体形态的动作监督数据规模化采集依然成本高昂。我们提出PointAction框架,通过显式的点基四维建模将视频预测与机器人行为相衔接。该框架对基础视频生成模型进行微调,使其联合预测未来RGB帧与动态三维点图,生成任务相关场景几何的时序一致三维运动。这些点动态作为结构化、具身无关的动作接口,由基于扩散的行为解码器映射为可执行机器人动作。通过将度量三维点动态作为视频预测与控制之间的桥梁,PointAction降低了仅依赖RGB的视觉动作锚定模糊性,支持在有限动作监督下实现跨任务与跨实体形态的迁移。实验表明,PointAction在机器人场景中达到了最先进的四维生成质量,在仿真环境中优于现有基准方法,并能泛化至预训练阶段未见过的两台真实机器人机械臂。