Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance. However, current HiL-RL frameworks remain intervention-intensive, relying on frequent human corrections to redirect the policy out of unproductive exploration, which incurs high labor cost and limits real-world scalability. To address this, we propose UniIntervene, an agentic intervention model that detects unproductive exploration and autonomously recovers the policy toward high-value states, taking over the bulk of interventions from human operators. Specifically, UniIntervene first performs future-conditioned action-value estimation, predicting the latent consequence of the current action and evaluating its induced value, which provides a more stable progress signal. Building on this, a temporal value-risk critic aggregates recent value dynamics and triggers intervention when the estimated value exhibits sustained stagnation or degradation. When intervention is required, UniIntervene retrieves a high-value recovery target from a memory of past intervention episodes and produces executable corrective actions through a goal-conditioned recovery policy. In this way, UniIntervene turns intervention from passive human correction into a value-aware recovery process for efficient real-world RL. Extensive experiments on diverse real-world manipulation tasks demonstrate that UniIntervene improves the average success rate by 8.6% while reducing human interventions by 57% relative to state-of-the-art HiL-RL baselines.


翻译:人机协同强化学习(HiL-RL)已成为真实世界机器人操作的有效范式,通过人类指导实现在线策略改进。然而,现有HiL-RL框架仍然依赖密集干预,需要频繁的人工纠偏以引导策略脱离无效探索,这导致高昂人力成本并限制真实场景扩展性。为此,我们提出UniIntervene智能体干预模型,该模型可检测无效探索并自主将策略恢复至高价值状态,从而承担人类操作员的大部分干预工作。具体而言,UniIntervene首先执行未来条件动作价值估计,预测当前动作的潜在后果并评估其诱导价值,从而提供更稳定的进程信号。基于此,时序价值风险评判器聚合近期价值动态,当估计价值呈现持续停滞或衰退时触发干预。当需要干预时,UniIntervene从过往干预事件记忆库中检索高价值恢复目标,并通过目标条件恢复策略生成可执行修正动作。通过这种方式,UniIntervene将干预从被动人工纠偏转变为面向高效真实世界强化学习的价值感知恢复过程。在多种真实世界操作任务上的大量实验表明,相较于现有最优HiL-RL基线方法,UniIntervene在将人类干预减少57%的同时将平均成功率提升8.6%。

0
下载
关闭预览

相关内容

《网络战仿真中的多智能体强化学习》最新42页报告
专知会员服务
48+阅读 · 2023年7月11日
【UIUC博士论文】高效多智能体深度强化学习,130页pdf
专知会员服务
76+阅读 · 2023年1月14日
【硬核书】迁移学习多智能体强化学习系统,131页pdf
专知会员服务
148+阅读 · 2022年7月8日
「基于通信的多智能体强化学习」 进展综述
【MIT博士论文】数据高效强化学习,176页pdf
【2022新书】强化学习工业应用
专知
18+阅读 · 2022年2月3日
强化学习精品书籍
平均机器
26+阅读 · 2019年1月2日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
4+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
8+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
8+阅读 · 7月31日
相关基金
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
23+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
19+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员