Generative models such as diffusion have been employed as world models in offline reinforcement learning to generate synthetic data for more effective learning. Existing work either generates diffusion models one-time prior to training or requires additional interaction data to update it. In this paper, we propose a novel approach for offline reinforcement learning with closed-loop policy evaluation and world-model adaptation. It iteratively leverages a guided diffusion world model to directly evaluate the offline target policy with actions drawn from it, and then performs an importance-sampled world model update to adaptively align the world model with the updated policy. We analyzed the performance of the proposed method and provided an upper bound on the return gap between our method and the real environment under an optimal policy. The result sheds light on various factors affecting learning performance. Evaluations in the D4RL environment show significant improvement over state-of-the-art baselines, especially when only random or medium-expertise demonstrations are available -- thus requiring improved alignment between the world model and offline policy evaluation.
翻译:扩散模型等生成模型已被用作离线强化学习中的世界模型,以生成合成数据来实现更有效的学习。现有工作要么在训练前一次性生成扩散模型,要么需要额外的交互数据来更新模型。本文提出了一种新颖的离线强化学习方法,该方法采用闭环策略评估与世界模型自适应机制。该方法迭代地利用引导扩散世界模型,通过从中采样的动作直接评估离线目标策略,随后执行重要性采样的世界模型更新,以自适应地将世界模型与更新后的策略对齐。我们分析了所提方法的性能,并给出了在最优策略下该方法与真实环境之间回报差距的上界。该结果揭示了影响学习性能的各种因素。在D4RL环境中的评估表明,相较于现有最先进的基线方法,本方法取得了显著改进,尤其是在仅有随机或中等专业水平演示可用的情况下——这要求世界模型与离线策略评估之间实现更好的对齐。