Diffusion-based visuomotor policies built on 3D visual representations have achieved strong performance in learning complex robotic skills. However, most existing methods employ an oversized denoising decoder. While increasing model capacity can improve denoising, empirical evidence suggests that it also introduces redundancy and noise in intermediate feature blocks. Crucially, we find that randomly masking backbone features in U-Net or skipping intermediate layers in DiT at inference time (without changing training) can improve performance, confirming the presence of task-irrelevant noise in intermediate features. To this end, we propose Variational Regularization (VR), a plug-and-play module that imposes a context-conditioned Gaussian over the noisy features and applies a KL-divergence regularizer, forming an adaptive information bottleneck. Extensive experiments on three simulation benchmarks, RoboTwin2.0, Adroit, and MetaWorld, show that our approach consistently improves task success rates over the baseline for both DP3-UNet and DP3-DiT, achieving new state-of-the-art results. Real-world experiments further demonstrate that our method performs well in practical deployments.
翻译:基于3D视觉表示的扩散视觉运动策略在学习复杂机器人技能方面已展现出强大性能。然而,现有方法大多采用过大的去噪解码器。虽然增加模型容量可提升去噪效果,但经验证据表明,这也会在中间特征块中引入冗余和噪声。关键的是,我们发现在推理时随机遮蔽U-Net主干特征或跳过DiT中间层(无需改变训练过程)可提升性能,这证实了中间特征中存在任务无关噪声。为此,我们提出变分正则化(VR)——一种即插即用模块,对带噪特征施加上下文条件高斯分布,并应用KL散度正则化器,形成自适应信息瓶颈。在RoboTwin2.0、Adroit和MetaWorld三个仿真基准上的大量实验表明,我们的方法在DP3-UNet和DP3-DiT上均能持续提升基线任务成功率,达到了新的最佳结果。真实世界实验进一步证明,该方法在实际部署中表现优异。