Real-time execution, enabled by asynchronous inference that ensures both smooth action trajectories and fast reactivity, is critical for realistic deployments of large-scale Vision-Language-Action models. However, recent work on real-time execution primarily focuses on variants of diffusion policies, even though it is more critical for autoregressive policies given their slower rollout speed in synchronous inference. In contrast, we demonstrate that autoregressive policies can achieve real-time execution by adjusting the tokenization horizon and applying constrained decoding, thereby guaranteeing strict latency bounds that enable multi-trajectory decoding to maximize performance. Across simulated and real-world environments, we find that the autoregressive policy consistently outperforms its equivalent-level flow-matching policy counterpart while achieving significantly improved task completion speeds from synchronous inference. Coupled with the inherent advantages of autoregressive policies, such as faster convergence and better generalizability in instruction-following, these results confirm that autoregressive policies can remain a competitive policy type supporting real-time execution.
翻译:实时执行通过异步推理实现平滑动作轨迹与快速响应能力,对于大规模视觉-语言-动作模型的实际部署至关重要。然而,近期关于实时执行的研究主要聚焦于扩散策略变体——尽管考虑到自回归策略在同步推理中较慢的 rollout 速度,其更需要实现实时执行。与此相反,我们证明通过调整分词视界(tokenization horizon)并应用约束解码,自回归策略能够实现实时执行,从而保证严格的延迟界限以支持多轨迹解码来最大化性能。在仿真与真实环境实验中,我们发现自回归策略在始终保持优于同等水平流匹配策略性能的同时,相较于同步推理显著提升了任务完成速度。结合自回归策略在指令跟随任务中收敛更快、泛化性更佳等固有优势,这些结果证实自回归策略仍可作为支持实时执行的竞争性策略类型。