Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and action execution. In real-time deployment, these models must process continuous visual streams, incurring substantial computational overhead. Visual token pruning -- a mainstream technique for accelerating Vision-Language Models (VLMs) by retaining salient tokens while discarding redundant ones -- offers a natural candidate solution to this challenge. However, directly applying VLM-oriented pruning methods to VLA inference can cause severe degradation in manipulation performance. Our analysis attributes this degradation to a key mismatch: VLA inference exhibits distinct attention patterns between the vision-language prefill stage and the action-decode stage, so pruning based only on context-prefill semantic salience is biased toward semantic cues and may remove action-critical visual tokens. Motivated by this observation, we propose VLA-Pruner, an effective plug-and-play token pruning method grounded in the visual requirements of VLA inference, further exploiting the temporal continuity of robot manipulation. Specifically, VLA-Pruner estimates visual-token importance from both semantic prefilling and temporally smoothed action relevance, and then applies a Combine-then-Filter strategy to retain compact, non-redundant tokens under the compute budget. Experiments show that VLA-Pruner outperforms state-of-the-art approaches across multiple VLA architectures, achieving up to 1.99x speedup with comparable manipulation quality.


翻译:视觉-语言-动作模型通过整合视觉感知、语言理解和动作执行,在具身人工智能领域展现出巨大潜力。在实时部署中,这些模型需处理连续视觉流,导致巨大计算开销。视觉标记剪枝——通过保留关键标记并丢弃冗余标记来加速视觉-语言模型的主流技术——为解决该挑战提供了天然候选方案。然而,将面向视觉语言模型的剪枝方法直接应用于视觉语言动作推理会导致操作性能严重下降。我们的分析将这种性能退化归因于关键不匹配:视觉语言动作推理在视觉语言预填充阶段与动作解码阶段表现出不同的注意力模式,因此仅基于上下文预填充语义显著性进行剪枝会偏向语义线索,并可能移除对动作关键的视觉标记。基于这一发现,我们提出VLA-Pruner——一种基于视觉语言动作推理视觉需求的有效即插即用标记剪枝方法,进一步利用机器人操作的时间连续性。具体而言,VLA-Pruner从语义预填充和时间平滑的动作相关性两方面估计视觉标记重要性,随后采用“先合并后过滤”策略在计算预算下保留紧凑且无冗余的标记。实验表明,VLA-Pruner在多种视觉语言动作架构上均优于现有方法,在保持相当操作质量的同时实现最高1.99倍加速。

0
下载
关闭预览

相关内容

视觉-语言-动作(VLA)模型的前世今生
专知会员服务
22+阅读 · 2025年8月29日
在回答之前先解释:组合视觉推理综述
专知会员服务
15+阅读 · 2025年8月27日
视觉语言动作模型:概念、进展、应用与挑战
专知会员服务
19+阅读 · 2025年5月18日
【博士论文】学习视觉-语言表示以实现多模态理解
专知会员服务
28+阅读 · 2025年2月8日
《深度神经网络剪枝》最新2023综述
专知会员服务
36+阅读 · 2023年8月17日
【博士论文】视觉语言交互中的视觉推理研究
专知会员服务
65+阅读 · 2021年12月1日
见微知著:语义分割中的弱监督学习
深度学习大讲堂
11+阅读 · 2017年12月6日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关VIP内容
相关资讯
见微知著:语义分割中的弱监督学习
深度学习大讲堂
11+阅读 · 2017年12月6日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员