Diffusion-based Vision-Language-Action (VLA) policies enable strong generalization in robotic manipulation, but remain sensitive to spurious visual correlations and noisy action generation, leading to brittle behavior under perturbations. We introduce Selected Diffusion Noise (SDN), a simple, training-free test-time method that improves both robustness and success rate by leveraging the diffusion noise space as a controllable degree of freedom. SDN dynamically samples noise vectors that are maximally separated from a reference set to mitigate reliance on spurious cues, while selecting candidates that yield more coherent action trajectories. This dual objective encourages stable behavior even under object-masked observations and reduces action jitter without modifying model parameters. We evaluate SDN on two simulation benchmarks (Google Robot, Widow-X) and two real-world robotic datasets across multiple VLA policies, including pi_0, Groot-N1.5, and Groot-N1.6. SDN consistently improves success rates by +8% in simulation and +10% in real-world settings, while producing smoother and more stable actions. Our results highlight that diffusion noise selection can serve as an effective and general mechanism for enhancing VLA policies at test time.
翻译:基于扩散的视觉-语言-动作(VLA)策略在机器人操作中展现出强大的泛化能力,但仍易受虚假视觉关联和噪声动作生成的影响,导致在扰动下表现脆弱。我们提出选择性扩散噪声(SDN),这是一种无需训练、即插即用的测试时方法,通过利用扩散噪声空间作为可控自由度来同时提升鲁棒性和任务成功率。SDN动态采样与参考集最大分离的噪声向量,以减轻对虚假线索的依赖,同时选择能产生更连贯动作轨迹的候选噪声。该双重目标机制在物体遮挡观测条件下仍能促进稳定行为,并在不修改模型参数的情况下减少动作抖动。我们在两个仿真基准(Google Robot、Widow-X)和两个真实机器人数据集上,基于多种VLA策略(包括pi_0、Groot-N1.5和Groot-N1.6)评估SDN。SDN在仿真中持续提升+8%的成功率,在真实场景中提升+10%,同时产生更平滑、更稳定的动作。我们的研究结果表明,扩散噪声选择可作为测试时增强VLA策略的有效通用机制。