Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imperfect rewards remains difficult due to inefficient exploration and poor credit assignment. Vision-Language-Action (VLA) models leverage large-scale multimodal pretraining to provide generalist, task-level reasoning, but current limitations hinder their direct use in fast and precise manipulation. In this paper, we propose Vision-Language-Action Jump-Starting (VLAJS), a method that bridges sparse VLA guidance with on-policy RL to improve exploration and learning efficiency. VLAJS treats VLAs as transient sources of high-level action suggestions that bias early exploration and improve credit assignment, while preserving the high-frequency, state-based control of RL. Our approach augments Proximal Policy Optimization (PPO) with a directional action-consistency regularization that softly aligns the RL agent's actions with VLA guidance during early training, without enforcing strict imitation, requiring demonstrations, or relying on continuous teacher queries. VLA guidance is applied sparsely and annealed over time, allowing the agent to adapt online and ultimately surpass the guiding policy. We evaluate VLAJS on six challenging manipulation tasks: lifting, pick-and-place, peg reorientation, peg insertion, poking, and pushing in simulation, and validate a subset on a real Franka Panda robot. VLAJS consistently outperforms PPO and distillation-style baselines in sample efficiency, reducing required environment interactions by over 50% in several tasks. Real-world experiments demonstrate zero-shot sim-to-real transfer and robust execution under clutter, object variation, and external perturbations.


翻译:强化学习(RL)为机器人操作提供了高频闭环控制能力,但由于探索效率低下和信用分配困难等原因,在稀疏奖励或不完美奖励的长时域任务中难以扩展。视觉-语言-动作(VLA)模型通过大规模多模态预训练实现了泛化型任务级推理能力,但当前局限性阻碍了其在快速精准操作任务中的直接应用。本文提出大语言-视觉-动作联合引导方法(VLAJS),该方法通过桥接稀疏VLA引导与在线策略RL,显著提升探索效率和学习性能。VLAJS将VLA作为瞬态高层动作建议源,在引导早期探索方向的同时改善信用分配,同时保留RL基于状态的高频控制特性。我们通过在近端策略优化(PPO)中引入方向性动作一致性正则化项,使得RL智能体在训练初期软对齐VLA引导动作,该方法既无需严格模仿、演示样本,也不依赖持续教师查询。VLA引导采用稀疏调度并随时间衰减退火,使智能体能够在线自适应并最终超越引导策略。我们在仿真环境中对六项挑战性操作任务(举升、抓取放置、销钉重定向、销钉插接、戳动、推挤)进行了评估,并在真实Franka Panda机器人上验证了部分任务。实验表明,VLAJS在样本效率上持续优于PPO及知识蒸馏基线,在多个任务中将所需环境交互次数降低50%以上。真实世界实验验证了零样本仿真到真机迁移能力,并在杂乱场景、物体形态变化及外部扰动下展现出鲁棒执行性能。

0
下载
关闭预览

相关内容

大语言模型智能体强化学习:全景综述
专知会员服务
51+阅读 · 2025年12月18日
面向大语言模型的智能体化强化学习图景:综述
专知会员服务
56+阅读 · 2025年9月3日
《机器人强化学习技术进展》34页
专知会员服务
40+阅读 · 2025年7月16日
大语言模型的强化学习技术综述
专知会员服务
42+阅读 · 2025年7月8日
《改进单智能体和多智能体深度强化学习方法》219页
专知会员服务
64+阅读 · 2025年2月14日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
「基于通信的多智能体强化学习」 进展综述
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
【MIT博士论文】数据高效强化学习,176页pdf
关于强化学习(附代码,练习和解答)
深度学习
38+阅读 · 2018年1月30日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
大语言模型智能体强化学习:全景综述
专知会员服务
51+阅读 · 2025年12月18日
面向大语言模型的智能体化强化学习图景:综述
专知会员服务
56+阅读 · 2025年9月3日
《机器人强化学习技术进展》34页
专知会员服务
40+阅读 · 2025年7月16日
大语言模型的强化学习技术综述
专知会员服务
42+阅读 · 2025年7月8日
《改进单智能体和多智能体深度强化学习方法》219页
专知会员服务
64+阅读 · 2025年2月14日
自动驾驶中的多智能体强化学习综述
专知会员服务
48+阅读 · 2024年8月20日
相关基金
国家自然科学基金
15+阅读 · 2016年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员