We show that offline actor-critic reinforcement learning can scale to large models - such as transformers - and follows similar scaling laws as supervised learning. We find that offline actor-critic algorithms can outperform strong, supervised, behavioral cloning baselines for multi-task training on a large dataset containing both sub-optimal and expert behavior on 132 continuous control tasks. We introduce a Perceiver-based actor-critic model and elucidate the key model features needed to make offline RL work with self- and cross-attention modules. Overall, we find that: i) simple offline actor critic algorithms are a natural choice for gradually moving away from the currently predominant paradigm of behavioral cloning, and ii) via offline RL it is possible to learn multi-task policies that master many domains simultaneously, including real robotics tasks, from sub-optimal demonstrations or self-generated data.
翻译:我们表明,离线演员-评论家强化学习能够扩展至大型模型(例如Transformer),并遵循与监督学习相似的缩放定律。研究发现,在一包含132个连续控制任务、同时涵盖次优与专家行为的大型数据集上进行多任务训练时,离线演员-评论家算法可超越强大的监督式行为克隆基线方法。我们提出了一种基于感知器(Perceiver)的演员-评论家模型,并阐明了使离线强化学习能够与自注意力及交叉注意力模块协同工作的关键模型特征。总体而言,我们发现:(i)简单的离线演员-评论家算法是逐步脱离当前主流行为克隆范式的自然选择;(ii)通过离线强化学习,有可能从次优示范或自生成数据中学习同时掌握多个领域(包括真实机器人任务)的多任务策略。