The utilization of broad datasets has proven to be crucial for generalization for a wide range of fields. However, how to effectively make use of diverse multi-task data for novel downstream tasks still remains a grand challenge in robotics. To tackle this challenge, we introduce a framework that acquires goal-conditioned policies for unseen temporally extended tasks via offline reinforcement learning on broad data, in combination with online fine-tuning guided by subgoals in learned lossy representation space. When faced with a novel task goal, the framework uses an affordance model to plan a sequence of lossy representations as subgoals that decomposes the original task into easier problems. Learned from the broad data, the lossy representation emphasizes task-relevant information about states and goals while abstracting away redundant contexts that hinder generalization. It thus enables subgoal planning for unseen tasks, provides a compact input to the policy, and facilitates reward shaping during fine-tuning. We show that our framework can be pre-trained on large-scale datasets of robot experiences from prior work and efficiently fine-tuned for novel tasks, entirely from visual inputs without any manual reward engineering.
翻译:广泛数据集的利用已被证明对多个领域的泛化至关重要。然而,如何有效利用多样化的多任务数据来处理新的下游任务,仍是机器人领域的一大挑战。为解决此挑战,我们提出一个框架,通过在广泛数据上进行离线强化学习,结合在学得的有损表示空间中由子目标引导的在线微调,来获取针对未见过的长时间延展任务的目标条件策略。面对新的任务目标时,该框架使用可供性模型规划一系列有损表示作为子目标,将原始任务分解为更易处理的问题。从广泛数据中学到的有损表示强调状态和目标中与任务相关的信息,同时抽象掉阻碍泛化的冗余上下文。从而它能够为未见任务进行子目标规划,为策略提供紧凑输入,并促进微调过程中的奖励塑形。我们展示了该框架可在先前工作的大规模机器人经验数据集上进行预训练,并高效微调以适应新任务,整个过程完全基于视觉输入且无需任何手动奖励工程。