Long-horizon robotic manipulation remains challenging for reinforcement learning (RL) because sparse rewards provide limited guidance for credit assignment. Practical policy improvement thus relies on richer intermediate supervision, such as dense progress rewards, which are costly to obtain and ill-suited to non-monotonic behaviors such as backtracking and recovery. To address this, we propose Advantage Reward Modeling (ARM), a framework that shifts from hard-to-quantify absolute progress to estimating relative advantage. We introduce a cost-effective tri-state labeling strategy -- Progressive, Regressive, and Stagnant -- that reduces human cognitive overhead while ensuring high cross-annotator consistency. By training on these intuitive signals, ARM enables automated progress annotation for both complete demonstrations and fragmented DAgger-style data. Integrating ARM into an offline RL pipeline allows for adaptive action-reward reweighting, effectively filtering suboptimal samples. Our approach achieves a 99.4% success rate on a challenging long-horizon towel-folding task, demonstrating improved stability and data efficiency over current VLA baselines with near-zero human intervention during policy training.
翻译:长时域机器人操控对强化学习仍具挑战性,因为稀疏奖励在信用分配上提供的指导有限。实用的策略改进依赖于更丰富的中间监督信号(如密集进度奖励),但这些信号获取成本高昂,且不适用于回溯、恢复等非单调行为。为此,我们提出优势奖励建模框架,将难以量化的绝对进度估计转变为相对优势估计。我们引入一种低成本的三态标注策略——进步、退化、停滞——在降低人类认知负荷的同时确保跨标注者一致性。通过基于这些直观信号进行训练,ARM可对完整演示和碎片化DAgger风格数据实现自动化进度标注。将ARM集成至离线强化学习流水线后,可实现自适应动作-奖励重新加权,有效过滤次优样本。在具有挑战性的长时域毛巾折叠任务中,本方法达到99.4%的成功率,在策略训练阶段几乎无需人工干预的情况下,比现有VLA基线展现出更优的稳定性与数据效率。