Representation learning and exploration are among the key challenges for any deep reinforcement learning agent. In this work, we provide a singular value decomposition based method that can be used to obtain representations that preserve the underlying transition structure in the domain. Perhaps interestingly, we show that these representations also capture the relative frequency of state visitations, thereby providing an estimate for pseudo-counts for free. To scale this decomposition method to large-scale domains, we provide an algorithm that never requires building the transition matrix, can make use of deep networks, and also permits mini-batch training. Further, we draw inspiration from predictive state representations and extend our decomposition method to partially observable environments. With experiments on multi-task settings with partially observable domains, we show that the proposed method can not only learn useful representation on DM-Lab-30 environments (that have inputs involving language instructions, pixel images, and rewards, among others) but it can also be effective at hard exploration tasks in DM-Hard-8 environments.
翻译:表示学习与探索是深度强化学习智能体面临的关键挑战。本研究提出一种基于奇异值分解的方法,用于获取保留领域内潜在转移结构的表示。有趣的是,我们证明这些表示同时捕捉了状态访问的相对频率,从而无需额外计算即可获得伪计数估计。为将该分解方法扩展至大规模领域,我们提出的算法无需构建转移矩阵,可借助深度网络实现,并支持小批量训练。此外,我们借鉴预测状态表示的思想,将分解方法扩展至部分可观测环境。通过在混合任务设置的部分可观测领域实验(涉及语言指令、像素图像和奖励等输入形式的DM-Lab-30环境),以及DM-Hard-8环境中的困难探索任务,我们证明该方法不仅能学习到有效表示,而且在硬探索任务中同样表现优异。