A long-standing goal of reinforcement learning is to acquire agents that can learn on training tasks and generalize well on unseen tasks that may share a similar dynamic but with different reward functions. The ability to generalize across tasks is important as it determines an agent's adaptability to real-world scenarios where reward mechanisms might vary. In this work, we first show that training a general world model can utilize similar structures in these tasks and help train more generalizable agents. Extending world models into the task generalization setting, we introduce a novel method named Task Aware Dreamer (TAD), which integrates reward-informed features to identify consistent latent characteristics across tasks. Within TAD, we compute the variational lower bound of sample data log-likelihood, which introduces a new term designed to differentiate tasks using their states, as the optimization objective of our reward-informed world models. To demonstrate the advantages of the reward-informed policy in TAD, we introduce a new metric called Task Distribution Relevance (TDR) which quantitatively measures the relevance of different tasks. For tasks exhibiting a high TDR, i.e., the tasks differ significantly, we illustrate that Markovian policies struggle to distinguish them, thus it is necessary to utilize reward-informed policies in TAD. Extensive experiments in both image-based and state-based tasks show that TAD can significantly improve the performance of handling different tasks simultaneously, especially for those with high TDR, and display a strong generalization ability to unseen tasks.
翻译:强化学习的长期目标之一是训练出能够在训练任务上学习,并在共享相似动力学但奖励函数不同的未见任务上实现良好泛化的智能体。跨任务泛化能力至关重要,因为它决定了智能体在奖励机制可能变化的真实场景中的适应性。本文首先证明,训练通用世界模型能够利用这些任务中的相似结构,从而帮助训练更具泛化能力的智能体。通过将世界模型扩展到任务泛化场景,我们提出了一种名为任务感知梦想家(TAD)的新方法,该方法整合了奖励信息特征,以识别跨任务的一致性潜在特征。在TAD中,我们计算样本数据对数似然的变分下界,引入一项新术语:通过状态区分任务,并将其作为奖励信息世界模型的优化目标。为展示TAD中奖励信息策略的优势,我们提出了一项新指标——任务分布相关性(TDR),用于定量衡量不同任务之间的相关性。对于高TDR任务(即任务差异显著的情况),我们证明马尔可夫策略难以区分它们,因此必须利用TAD中的奖励信息策略。在基于图像和基于状态的任务上的大量实验表明,TAD能够显著提升同时处理不同任务的性能,尤其对于高TDR任务,并在未见任务上展现出强大的泛化能力。