Sparse rewards are one of the factors leading to low sample efficiency in multi-goal reinforcement learning (RL). Based on Hindsight Experience Replay (HER), model-based relabeling methods have been proposed to relabel goals using virtual trajectories obtained by interacting with the trained model, which can effectively enhance the sample efficiency in accurately modelable sparse-reward environments. However, they are ineffective in robot manipulation environment. In our paper, we design a robust framework called Robust Model-based Hindsight Experience Replay (RoMo-HER) which can effectively utilize the dynamical model in robot manipulation environments to enhance the sample efficiency. RoMo-HER is built upon a dynamics model and a novel goal relabeling technique called Foresight relabeling (FR), which selects the prediction starting state with a specific strategy, predicts the future trajectory of the starting state, and then relabels the goal using the dynamics model and the latest policy to train the agent. Experimental results show that RoMo-HER has higher sample efficiency than HER and Model-based Hindsight Experience Replay in several simulated robot manipulation environments. Furthermore, we integrate RoMo-HER and Relay Hindsight Experience Replay (RHER), which currently exhibits the highest sampling efficiency in most benchmark environments, resulting in a novel approach called Robust Model-based Relay Hindsight Experience Replay (RoMo-RHER). Our experimental results demonstrate that RoMo-RHER achieves higher sample efficiency over RHER, outperforming RHER by 25% and 26% in FetchPush-v1 and FetchPickandPlace-v1, respectively.
翻译:稀疏奖励是导致多目标强化学习(RL)样本效率低下的因素之一。基于回溯经验回放(HER),学者们提出了基于模型的重标记方法,通过与训练模型交互获得的虚拟轨迹来重标记目标,从而在可精确建模的稀疏奖励环境中有效提升样本效率。然而,这些方法在机器人操作环境中效果不佳。本文设计了一种名为鲁棒基于模型回溯经验回放(RoMo-HER)的鲁棒框架,该框架能够有效利用机器人操作环境中的动力学模型来提升样本效率。RoMo-HER基于动力学模型和一种名为前瞻重标记(FR)的新型目标重标记技术构建,该技术采用特定策略选择预测起始状态,预测起始状态的未来轨迹,然后利用动力学模型和最新策略重标记目标,用于训练智能体。实验结果表明,在多个模拟机器人操作环境中,RoMo-HER的样本效率高于HER和基于模型的回溯经验回放。此外,我们将RoMo-HER与目前大多数基准环境中具有最高采样效率的中继回溯经验回放(RHER)相结合,提出了一种名为鲁棒基于模型的中继回溯经验回放(RoMo-RHER)的新方法。实验结果表明,RoMo-RHER相比RHER实现了更高的样本效率,在FetchPush-v1和FetchPickandPlace-v1环境中分别超出RHER 25%和26%。