Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA policies due to high-dimensional hand control and compounding execution errors, which makes real-world RL post-training essential for bridging the gap between visually grounded action generation and physically reliable dexterous execution. However, high-dimensional dexterous exploration often triggers temporal inconsistency, sample inefficiency and hardware risks in the real world. To address these challenges, we propose BORA, an offline-to-online RL post-training framework designed for real-world dexterous VLA models. In the offline phase, BORA constructs a critic that takes both the VLM's cognition tokens and action chunks as inputs. This design enables action-conditioned value guidance, allowing the critic to evaluate dexterous hand motions beyond visual context alone. During the subsequent online phase, BORA freezes the VLA base and introduces a lightweight, Human-in-the-Loop (HiL) chunk-wise residual adaptation mechanism to mitigate real-world execution errors and further correct the offline-learned intents within the actual physical environment. By inheriting the offline critic and employing intervention-driven rewards, BORA effectively corrects execution discrepancies and adapts to real-world physical variances while preserving the pretrained policy as a stable prior. Extensive evaluations across five complex real-world dexterous tasks demonstrate that BORA significantly outperforms pure imitation learning and traditional decoupled RL baselines, achieving a 33% absolute increase in average success rate under standard settings and up to a 43% improvement in unseen object generalization.
翻译:摘要:视觉-语言-动作(VLA)模型已成为将视觉-语言理解映射到真实世界机器人操作中的一种有前景的范式。然而,由于灵巧操作涉及高维手部控制及累积执行误差,VLA策略仍面临挑战,这使得现实环境中的强化学习后训练成为弥合视觉驱动动作生成与物理可靠灵巧执行之间鸿沟的关键。但高维灵巧探索常引发时间不一致性、样本低效及硬件风险。为此,我们提出BORA——一种面向真实世界灵巧VLA模型的离策略到在策略强化学习后训练框架。在离线阶段,BORA构建了一个同时以VLM认知标记和动作块为输入的评论家。该设计实现了动作条件化价值引导,使评论家能够超越单纯视觉上下文评估灵巧手部运动。在随后的在策略阶段,BORA冻结VLA基底,引入轻量级人在环路式块级残差自适应机制,以缓解真实世界执行误差,并在实际物理环境中进一步校正离线习得意图。通过继承离线评论家并采用干预驱动奖励,BORA在保留预训练策略作为稳定先验的同时,有效校正执行偏差并适应真实物理差异。在五项复杂真实世界灵巧任务上的广泛评估表明,BORA显著优于纯模仿学习及传统解耦强化学习基线,在标准设置下平均成功率提升33%,在未见物体泛化中提升达43%。