Due to the trial-and-error nature, it is typically challenging to apply RL algorithms to safety-critical real-world applications, such as autonomous driving, human-robot interaction, robot manipulation, etc, where such errors are not tolerable. Recently, safe RL (i.e. constrained RL) has emerged rapidly in the literature, in which the agents explore the environment while satisfying constraints. Due to the diversity of algorithms and tasks, it remains difficult to compare existing safe RL algorithms. To fill that gap, we introduce GUARD, a Generalized Unified SAfe Reinforcement Learning Development Benchmark. GUARD has several advantages compared to existing benchmarks. First, GUARD is a generalized benchmark with a wide variety of RL agents, tasks, and safety constraint specifications. Second, GUARD comprehensively covers state-of-the-art safe RL algorithms with self-contained implementations. Third, GUARD is highly customizable in tasks and algorithms. We present a comparison of state-of-the-art safe RL algorithms in various task settings using GUARD and establish baselines that future work can build on.
翻译:由于强化学习具有试错特性,将其应用于自动驾驶、人机交互、机器人操作等安全关键型实际应用场景通常面临巨大挑战——这些场景中任何错误都不可容忍。近年来,安全强化学习(即约束强化学习)领域迅速发展,智能体在满足约束条件的同时探索环境。由于算法与任务的多样性,现有安全强化学习算法之间的比较仍然困难。为填补这一空白,我们提出了GUARD——一个通用统一的安全强化学习开发基准。与现有基准相比,GUARD具备以下优势:第一,GUARD是一个通用型基准,包含多样化的强化学习智能体、任务类型及安全约束规范;第二,GUARD全面覆盖了当前最先进的安全强化学习算法,并提供自包含实现;第三,GUARD在任务与算法层面均具有高度可定制性。我们使用GUARD对多种任务场景下的最先进安全强化学习算法进行对比,建立了可供后续研究参考的基线结果。