Robotic manipulation requires reasoning about future spatial-temporal interactions and geometric constraints, yet existing Vision-Language-Action (VLA) policies often leave predictive representation weakly coupled with action execution, causing failures in tasks requiring precise spatial-temporal coordination. We propose STARRY, a world-model-enhanced action-generation policy that aligns spatial-temporal prediction and action generation by jointly denoising future spatial-temporal latents and actions through a unified diffusion process. To bridge 2D visual tokens and 3D metric control, STARRY introduces Geometry-Aware Selective Attention Modulation (GASAM), which converts predicted depth and end-effector geometry into token-aligned weights for selective action-attention modulation. On RoboTwin 2.0, STARRY achieves 93.82% / 93.30% average success under Clean and Randomized settings across 50 bimanual tasks. Real-world experiments show that STARRY improves average success from 42.5% to 70.8% compared with $π_{0.5}$. These results demonstrate the effectiveness of action-centric spatial-temporal world modeling for spatially and temporally demanding robotic manipulation.
翻译:机器人操作需要推理未来的空间-时间交互与几何约束,然而现有的视觉-语言-动作(VLA)策略往往使预测表示与动作执行弱耦合,导致在需要精确空间-时间协调的任务中失败。本文提出STARRY,一种世界模型增强的动作生成策略,通过统一扩散过程对未来的空间-时间潜变量和动作进行联合去噪,从而对齐空间-时间预测与动作生成。为桥接2D视觉标记与3D度量控制,STARRY引入几何感知选择性注意力调制(GASAM),将预测深度与末端执行器几何信息转换为标记对齐的权重,用于选择性动作注意力调制。在RoboTwin 2.0上,STARRY在清洁和随机设置下,针对50个双臂任务实现了平均93.82%和93.30%的成功率。真实世界实验表明,与π₀.₅相比,STARRY将平均成功率从42.5%提升至70.8%。这些结果证明了面向动作的空间-时间世界建模在空间和时间要求严苛的机器人操作中的有效性。