Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show promising results in simple environments but often get stuck in environments with multimodal and stochastic dynamics. In this work, we propose a variational dynamic model based on the conditional variational inference to model the multimodality and stochasticity. We consider the environmental state-action transition as a conditional generative process by generating the next-state prediction under the condition of the current state, action, and latent variable, which provides a better understanding of the dynamics and leads a better performance in exploration. We derive an upper bound of the negative log-likelihood of the environmental transition and use such an upper bound as the intrinsic reward for exploration, which allows the agent to learn skills by self-supervised exploration without observing extrinsic rewards. We evaluate the proposed method on several image-based simulation tasks and a real robotic manipulating task. Our method outperforms several state-of-the-art environment model-based exploration approaches.
翻译:高效探索仍是强化学习中的挑战性问题,尤其当环境提供的外在奖励稀疏甚至完全缺失时。基于内在动机的显著进展在简单环境中展现出令人鼓舞的结果,但在具有多模态和随机动态特性的环境中常陷入困境。本文提出一种基于条件变分推断的变分动力学模型,用于建模多模态性和随机性。我们将环境状态-动作转移视为条件生成过程,通过在当前状态、动作与潜变量条件下生成下一状态预测,从而更深入理解动态特性并提升探索性能。我们推导出环境转移负对数似然的上界,并将其作为探索的内在奖励,使智能体无需外在奖励即可通过自监督探索学习技能。我们在多个基于图像的仿真任务及真实机器人操作任务上评估所提方法,结果表明其性能优于当前最先进的基于环境模型的探索方法。