In recent years, data-driven reinforcement learning (RL), also known as offline RL, have gained significant attention. However, the role of data sampling techniques in offline RL has been overlooked despite its potential to enhance online RL performance. Recent research suggests applying sampling techniques directly to state-transitions does not consistently improve performance in offline RL. Therefore, in this study, we propose a memory technique, (Prioritized) Trajectory Replay (TR/PTR), which extends the sampling perspective to trajectories for more comprehensive information extraction from limited data. TR enhances learning efficiency by backward sampling of trajectories that optimizes the use of subsequent state information. Building on TR, we build the weighted critic target to avoid sampling unseen actions in offline training, and Prioritized Trajectory Replay (PTR) that enables more efficient trajectory sampling, prioritized by various trajectory priority metrics. We demonstrate the benefits of integrating TR and PTR with existing offline RL algorithms on D4RL. In summary, our research emphasizes the significance of trajectory-based data sampling techniques in enhancing the efficiency and performance of offline RL algorithms.
翻译:近年来,数据驱动强化学习(即离线强化学习)受到了广泛关注。然而,尽管数据采样技术有潜力提升在线强化学习性能,其在离线强化学习中的作用却常被忽视。近期研究表明,直接将采样技术应用于状态转移并不能持续改善离线强化学习的性能。为此,本文提出一种记忆技术——(优先)轨迹回放(TR/PTR),将采样视角扩展至轨迹层面,以从有限数据中提取更全面的信息。TR通过反向采样轨迹优化后续状态信息的利用,从而提升学习效率。在TR基础上,我们构建了加权评论家目标以避免离线训练中采样未见动作,并提出了优先轨迹回放(PTR),该技术通过多种轨迹优先级指标实现更高效的轨迹采样。我们在D4RL数据集上证明了将TR和PTR与现有离线强化学习算法集成的优势。总之,本研究强调了基于轨迹的数据采样技术在提升离线强化学习算法效率与性能中的重要性。