Pre-training large models on vast amounts of web data has proven to be an effective approach for obtaining powerful, general models in several domains, including language and vision. However, this paradigm has not yet taken hold in deep reinforcement learning (RL). This gap is due to the fact that the most abundant form of embodied behavioral data on the web consists of videos, which do not include the action labels required by existing methods for training policies from offline data. We introduce Latent Action Policies from Observation (LAPO), a method to infer latent actions and, consequently, latent-action policies purely from action-free demonstrations. Our experiments on challenging procedurally-generated environments show that LAPO can act as an effective pre-training method to obtain RL policies that can then be rapidly fine-tuned to expert-level performance. Our approach serves as a key stepping stone to enabling the pre-training of powerful, generalist RL models on the vast amounts of action-free demonstrations readily available on the web.
翻译:在大规模网络数据上预训练大型模型已被证明是在语言和视觉等多个领域获得强大通用模型的有效方法。然而,这一范式尚未在深度强化学习(RL)领域得到广泛应用。这种差距源于网络上最丰富的具身行为数据形式是视频,而视频不包含现有方法从离线数据训练策略所需的动作标签。我们提出基于观察的潜在动作策略(LAPO),一种纯粹从无动作示范中推断潜在动作并进而学习潜在动作策略的方法。在具有挑战性的程序化生成环境中的实验表明,LAPO可作为一种有效的预训练方法,获得能快速微调至专家级性能的强化学习策略。我们的方法为在互联网上大量可获取的无动作示范数据上预训练强大的通用强化学习模型奠定了关键基础。