Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models. However, enabling seamless humanoid interventions is a formidable systems challenge due to complex whole-body kinematics and dexterous-hand control. Consequently, the collected intervention trajectories are often suboptimal, and methods that rely on human interventions as expert supervision can absorb hesitant, inefficient, or even erroneous behaviors. To address both the system and algorithmic challenges, we propose ROVE, a reinforcement learning framework for humanoid VLA post-training with imperfect human interventions. First, ROVE introduces a human-in-the-loop pipeline capable of collecting deployment and intervention data for humanoid manipulation. Second, it utilizes Optimistic Value Estimation (OVE) to prioritize high-value behaviors from mixed-quality trajectories. To further robustify value estimation, we incorporate cross-embodiment human experience videos to provide rich supervision for long-tailed failure and recovery modes. The resulting critic yields informative advantage signals, steering the VLA actor to focus on high-value behaviors rather than indiscriminately imitating all actions. On challenging real-world contact-rich and fine-grained humanoid manipulation tasks, ROVE outperforms experience-learning baselines and consistently improves across multiple rollout-intervention iterations.
翻译:人类干预为视觉-语言-动作(VLA)模型的后训练提供了关键的纠正信号。然而,由于复杂全身运动学与灵巧手控制的挑战,实现无缝的人形干预是一项艰巨的系统性难题。因此,收集到的干预轨迹往往非最优,而依赖人类干预作为专家监督的方法可能吸收犹豫、低效甚至错误的动作。为同时解决系统与算法挑战,我们提出ROVE——一种利用非完美人类干预进行人形VLA后训练的强化学习框架。首先,ROVE引入人机闭环流水线,可采集人形机器人操作中的部署与干预数据。其次,它采用乐观值估计(OVE)从混合质量轨迹中优先提取高价值行为。为增强值估计的鲁棒性,我们融合跨本体的类人经验视频,为长尾故障与恢复模式提供丰富监督信号。由此得到的评价器可生成富有信息量的优势信号,引导VLA执行器聚焦高价值行为而非盲目模仿所有动作。在极具挑战性的真实世界接触密集与精细人形操作任务中,ROVE优于基于经验学习的基线方法,并在多轮部署-干预迭代中持续提升性能。