Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference methods primarily optimize trajectory smoothness, but neglect the critical latency in reacting to environmental changes. By rethinking the notion of reaction in action chunking policies, this paper presents a systematic analysis of the factors governing reaction time. We show that reaction time follows a uniform distribution determined jointly by the Time to First Action (TTFA) and the execution horizon. Moreover, we reveal that the standard practice of applying a constant schedule in flow-based VLAs can be inefficient and forces the system to complete all sampling steps before any movement can start, forming the bottleneck in reaction latency. To overcome this issue, we propose Fast Action Sampling for ImmediaTE Reaction (FASTER). By introducing a Horizon-Aware Schedule, FASTER adaptively prioritizes near-term actions during flow sampling, compressing the denoising of the immediate reaction by tenfold (e.g., in $π_{0.5}$ and X-VLA) into a single step, while preserving the quality of long-horizon trajectory. Coupled with a streaming client-server pipeline, FASTER substantially reduces the effective reaction latency on real robots, especially when deployed on consumer-grade GPUs. Real-world experiments, including a highly dynamic table tennis task, prove that FASTER unlocks unprecedented real-time responsiveness for generalist policies, enabling rapid generation of accurate and smooth trajectories.


翻译:实时执行对于在物理世界中部署视觉-语言-动作(VLA)模型至关重要。现有异步推理方法主要优化轨迹平滑性,但忽视了响应环境变化的关键延迟。通过重新审视动作分块策略中的反应机制,本文系统分析了控制反应时间的因素。我们证明反应时间服从由首次动作时间(TTFA)和执行周期共同决定的均匀分布。此外,我们发现流式VLA中采用恒定时间表的常规做法效率低下,迫使系统在开始任何动作前完成所有采样步骤,这构成了反应延迟的瓶颈。为解决此问题,我们提出即时反应的快速动作采样(FASTER)。通过引入周期感知时间表,FASTER在流采样过程中自适应地优先处理近期动作,将即时反应的去噪过程压缩十倍(例如在$π_{0.5}$和X-VLA中)为单步操作,同时保持长周期轨迹的质量。结合流式客户端-服务器流水线,FASTER显著降低了真实机器人上的有效反应延迟,尤其在消费级GPU上部署时。包含高动态乒乓球任务在内的真实世界实验证明,FASTER实现了通用策略前所未有的实时响应能力,能够快速生成准确且平滑的轨迹。

0
下载
关闭预览

相关内容

【综述】 机器人学习中的世界模型:全面综述
专知会员服务
21+阅读 · 5月4日
视觉-语言-动作(VLA)模型的前世今生
专知会员服务
21+阅读 · 2025年8月29日
【Flink】基于 Flink 的流式数据实时去重
AINLP
14+阅读 · 2020年9月29日
复现 | FastDVDNet:实时视频去噪算法
CVer
13+阅读 · 2019年7月12日
Fast-OCNet: 更快更好的OCNet.
极市平台
21+阅读 · 2019年2月10日
一文读懂目标检测:R-CNN、Fast R-CNN、Faster R-CNN、YOLO、SSD
七月在线实验室
11+阅读 · 2018年7月18日
ETP:精确时序动作定位
极市平台
13+阅读 · 2018年5月25日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 4月29日
VIP会员
相关主题
最新内容
边缘计算的军事应用
专知会员服务
2+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
5+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
12+阅读 · 8月7日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员