Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent studies have introduced tactile or force feedback into VLAs to address contact-rich tasks. However, these models are typically deployed as offline policies. When contact conditions shift from the training distribution, the policy cannot perform online adaptation, leading to problems such as inappropriate contact forces and inefficient retries. Therefore, we propose TORL-VLA, a tactile-guided online reinforcement learning framework that couples tactile feedback with policy refinement for contact-rich manipulation. Our method introduces a tactile-derived wrench-aware VLA to predict reference actions and future wrench sequences, while a lightweight online RL module is used to refine the reference actions. To stabilize learning from mixed exploratory policy-generated and human-intervention data, we introduce an intervention-censored critic that prevents post-intervention success from being wrongly credited to policy-generated actions preceding intervention. Real-robot experiments on long-horizon contact-rich tasks, including latch manipulation, coffee-cup placement, and egg handling, show that TORL-VLA improves success rates at both subtask and full-task levels, as well as time-bounded execution efficiency over strong baselines.
翻译:视觉-语言-动作(VLA)模型已成为机器人操作任务的强大框架,近期研究通过引入触觉或力反馈拓展VLA以处理密集接触任务。然而,这些模型通常作为离线策略部署。当接触条件偏离训练分布时,策略无法进行在线自适应,导致接触力不当、重试效率低下等问题。为此,我们提出TORL-VLA,一种面向密集接触操作的触觉引导在线强化学习框架,通过耦合触觉反馈与策略优化来解决上述问题。该方法引入基于触觉的力矩感知VLA预测参考动作与未来力矩序列,并采用轻量级在线强化学习模块优化参考动作。为稳定混合探索策略生成数据与人工干预数据的联合学习过程,我们提出干预审查批评者,防止干预后成功被错误归因于干预前的策略生成动作。在包含门闩操作、咖啡杯放置和鸡蛋处理的长时序密集接触任务中,真实机器人实验表明,TORL-VLA在子任务与全任务层面均能提升成功率,并在时限执行效率上超越现有强基线方法。