A core ambition of reinforcement learning (RL) is the creation of agents capable of rapid learning in novel tasks. Meta-RL aims to achieve this by directly learning such agents. Black box methods do so by training off-the-shelf sequence models end-to-end. By contrast, task inference methods explicitly infer a posterior distribution over the unknown task, typically using distinct objectives and sequence models designed to enable task inference. Recent work has shown that task inference methods are not necessary for strong performance. However, it remains unclear whether task inference sequence models are beneficial even when task inference objectives are not. In this paper, we present strong evidence that task inference sequence models are still beneficial. In particular, we investigate sequence models with permutation invariant aggregation, which exploit the fact that, due to the Markov property, the task posterior does not depend on the order of data. We empirically confirm the advantage of permutation invariant sequence models without the use of task inference objectives. However, we also find, surprisingly, that there are multiple conditions under which permutation variance remains useful. Therefore, we propose SplAgger, which uses both permutation variant and invariant components to achieve the best of both worlds, outperforming all baselines on continuous control and memory environments.
翻译:强化学习(RL)的核心目标之一是创建能够在新型任务中快速学习的智能体。元强化学习旨在通过直接学习此类智能体来实现这一目标。黑箱方法通过端到端地训练现成的序列模型来实现这一点。相比之下,任务推断方法则显式地推断未知任务的后验分布,通常使用专门设计用于任务推断的不同目标和序列模型。最新研究表明,任务推断方法并非获得强性能的必要条件。然而,即使在不使用任务推断目标的情况下,任务推断序列模型是否仍具有优势尚不明确。本文提供了强有力的证据表明,任务推断序列模型仍然具有优势。具体而言,我们研究了具有排列不变聚合的序列模型,该模型利用马尔可夫特性,即任务后验不依赖于数据的顺序。我们通过实验证实了在不使用任务推断目标的情况下排列不变序列模型的优势。然而,我们也意外地发现,存在多种条件下排列可变性仍然有效。因此,我们提出了SplAgger,该方法同时使用排列可变和排列不变组件以实现两者的最佳效果,在连续控制和记忆环境中全面超越所有基线方法。