Addressing decision-making problems using sequence modeling to predict future trajectories shows promising results in recent years. In this paper, we take a step further to leverage the sequence predictive method in wider areas such as long-term planning, vision-based control, and multi-task decision-making. To this end, we propose a method to utilize a diffusion-based generative sequence model to plan a series of milestones in a latent space and to have an agent to follow the milestones to accomplish a given task. The proposed method can learn control-relevant, low-dimensional latent representations of milestones, which makes it possible to efficiently perform long-term planning and vision-based control. Furthermore, our approach exploits generation flexibility of the diffusion model, which makes it possible to plan diverse trajectories for multi-task decision-making. We demonstrate the proposed method across offline reinforcement learning (RL) benchmarks and an visual manipulation environment. The results show that our approach outperforms offline RL methods in solving long-horizon, sparse-reward tasks and multi-task problems, while also achieving the state-of-the-art performance on the most challenging vision-based manipulation benchmark.
翻译:近年来,利用序列建模预测未来轨迹以解决决策问题的方法取得了显著进展。本文进一步将序列预测方法推广至长期规划、视觉控制及多任务决策等更广泛领域。为此,我们提出一种方法:采用基于扩散模型的生成式序列建模,在潜在空间中规划一系列里程碑,并通过智能体跟随这些里程碑完成给定任务。该方法可学习与任务控制相关且低维的里程碑潜在表征,从而实现高效的长期规划与视觉控制。此外,我们的方法利用扩散模型的生成灵活性,可规划多样化的轨迹以支持多任务决策。我们在离线强化学习基准测试和视觉操作环境中验证了所提方法。结果表明,在解决长时域、稀疏奖励任务及多任务问题时,本方法优于现有离线强化学习方法,并在最具挑战性的视觉操作基准中达到了最先进性能。