Episodic training, where an agent's environment is reset after every success or failure, is the de facto standard when training embodied reinforcement learning (RL) agents. The underlying assumption that the environment can be easily reset is limiting both practically, as resets generally require human effort in the real world and can be computationally expensive in simulation, and philosophically, as we'd expect intelligent agents to be able to continuously learn without intervention. Work in learning without any resets, i.e{.} Reset-Free RL (RF-RL), is promising but is plagued by the problem of irreversible transitions (e.g{.} an object breaking) which halt learning. Moreover, the limited state diversity and instrument setup encountered during RF-RL means that works studying RF-RL largely do not require their models to generalize to new environments. In this work, we instead look to minimize, rather than completely eliminate, resets while building visual agents that can meaningfully generalize. As studying generalization has previously not been a focus of benchmarks designed for RF-RL, we propose a new Stretch Pick-and-Place benchmark designed for evaluating generalizations across goals, cosmetic variations, and structural changes. Moreover, towards building performant reset-minimizing RL agents, we propose unsupervised metrics to detect irreversible transitions and a single-policy training mechanism to enable generalization. Our proposed approach significantly outperforms prior episodic, reset-free, and reset-minimizing approaches achieving higher success rates with fewer resets in Stretch-P\&P and another popular RF-RL benchmark. Finally, we find that our proposed approach can dramatically reduce the number of resets required for training other embodied tasks, in particular for RoboTHOR ObjectNav we obtain higher success rates than episodic approaches using 99.97\% fewer resets.
翻译:分段式训练(即在每次成功或失败后重置智能体环境)是训练具身强化学习智能体的标准范式。环境易于重置这一潜在假设既存在实际局限性(现实中重置通常需要人工操作,仿真中也可能产生高昂计算成本),也存在哲学缺陷(我们期望智能体无需干预即可持续学习)。零重置学习(即无重置强化学习)虽前景广阔,却受困于不可逆状态转换(如物体损坏)导致学习停滞的问题。此外,无重置强化学习过程中有限的状态多样性与仪器配置意味着相关研究大多无需模型在新环境中泛化。本研究旨在最小化而非完全消除重置,同时构建具备有效泛化能力的视觉智能体。鉴于现有无重置强化学习基准测试未聚焦泛化研究,我们提出新型Stretch拾取-放置基准,用于评估跨目标、外观变异及结构变化的泛化能力。为构建高性能最小化重置强化学习智能体,我们提出检测不可逆状态转换的无监督指标及支持泛化的单策略训练机制。所提方法在Stretch-P&P及另一主流无重置强化学习基准中显著优于现有分段式、无重置及最小化重置方法,以更少重置次数实现更高成功率。最终发现,该方法可显著减少其他具身任务训练所需重置次数,尤其是在RoboTHOR对象导航任务中,以降低99.97%重置次数的优势获得比分段式方法更高的成功率。