Meta Reinforcement Learning (Meta-RL) has seen substantial advancements recently. In particular, off-policy methods were developed to improve the data efficiency of Meta-RL techniques. \textit{Probabilistic embeddings for actor-critic RL} (PEARL) is a leading approach for multi-MDP adaptation problems. A major drawback of many existing Meta-RL methods, including PEARL, is that they do not explicitly consider the safety of the prior policy when it is exposed to a new task for the first time. Safety is essential for many real-world applications, including field robots and Autonomous Vehicles (AVs). In this paper, we develop the PEARL PLUS (PEARL$^+$) algorithm, which optimizes the policy for both prior (pre-adaptation) safety and posterior (after-adaptation) performance. Building on top of PEARL, our proposed PEARL$^+$ algorithm introduces a prior regularization term in the reward function and a new Q-network for recovering the state-action value under prior context assumptions, to improve the robustness to task distribution shift and safety of the trained network exposed to a new task for the first time. The performance of PEARL$^+$ is validated by solving three safety-critical problems related to robots and AVs, including two MuJoCo benchmark problems. From the simulation experiments, we show that safety of the prior policy is significantly improved and more robust to task distribution shift compared to PEARL.
翻译:元强化学习(Meta-RL)近年来取得了显著进展。特别是,离策略方法被开发用于提升元强化学习技术的数据效率。用于演员-评论家强化学习的概率嵌入(PEARL)是解决多MDP适应问题的领先方法。许多现有元强化学习方法(包括PEARL)的一个主要缺陷是,当先验策略首次暴露于新任务时,它们未明确考虑其安全性。安全性对于许多实际应用至关重要,包括野外机器人和自动驾驶汽车(AVs)。本文提出了PEARL PLUS(PEARL$^+$)算法,该算法同时优化先验(预适应)安全性与后验(适应后)性能的策略。基于PEARL,我们提出的PEARL$^+$算法在奖励函数中引入先验正则化项,并新增一个Q网络用于在先验上下文假设下恢复状态-动作值,从而提升训练网络首次暴露于新任务时对任务分布偏移的鲁棒性和安全性。通过解决三个与机器人和自动驾驶汽车相关的安全关键问题(包括两个MuJoCo基准问题),验证了PEARL$^+$的性能。仿真实验表明,与PEARL相比,先验策略的安全性显著提升,且对任务分布偏移具有更强的鲁棒性。