Pretraining RL models on offline video datasets is a promising way to improve their training efficiency in online tasks, but challenging due to the inherent mismatch in tasks, dynamics, and behaviors across domains. A recent model, APV, sidesteps the accompanied action records in offline datasets and instead focuses on pretraining a task-irrelevant, action-free world model within the source domains. We present Vid2Act, a model-based RL method that learns to transfer valuable action-conditioned dynamics and potentially useful action demonstrations from offline to online settings. The main idea is to use the world models not only as simulators for behavior learning but also as tools to measure the domain relevance for both dynamics representation transfer and policy transfer. Specifically, we train the world models to generate a set of time-varying task similarities using a domain-selective knowledge distillation loss. These similarities serve two purposes: (i) adaptively transferring the most useful source knowledge to facilitate dynamics learning, and (ii) learning to replay the most relevant source actions to guide the target policy. We demonstrate the advantages of Vid2Act over the action-free visual RL pretraining method in both Meta-World and DeepMind Control Suite.
翻译:在离线视频数据集上预训练强化学习模型是提升其在线任务训练效率的有效途径,但由于跨域任务、动态规律和行为模式的固有差异,这一方向面临挑战。近期提出的APV模型回避了离线数据集中伴随的动作记录,转而专注于在源域中预训练与任务无关、无动作条件的世界模型。我们提出Vid2Act——一种基于模型的强化学习方法,能够将源域中具有价值的动作条件动力学和潜在有用的动作示范迁移到在线设置中。核心思想在于:将世界模型不仅用作行为学习的模拟器,更作为衡量域相关性的工具,以同时实现动力学表征迁移和策略迁移。具体而言,我们通过领域选择性知识蒸馏损失函数训练世界模型生成一组时变任务相似度。这些相似度服务于两个目标:(i) 自适应迁移最有效的源域知识以促进动力学学习,(ii) 学习重放最相关的源域动作以引导目标策略。我们在Meta-World和DeepMind控制套件中验证了Vid2Act相较于无动作视觉强化学习预训练方法的优势。