Extracting a stable and compact representation of the environment is crucial for efficient reinforcement learning in high-dimensional, noisy, and non-stationary environments. Different categories of information coexist in such environments -- how to effectively extract and disentangle these information remains a challenging problem. In this paper, we propose IFactor, a general framework to model four distinct categories of latent state variables that capture various aspects of information within the RL system, based on their interactions with actions and rewards. Our analysis establishes block-wise identifiability of these latent variables, which not only provides a stable and compact representation but also discloses that all reward-relevant factors are significant for policy learning. We further present a practical approach to learning the world model with identifiable blocks, ensuring the removal of redundants but retaining minimal and sufficient information for policy optimization. Experiments in synthetic worlds demonstrate that our method accurately identifies the ground-truth latent variables, substantiating our theoretical findings. Moreover, experiments in variants of the DeepMind Control Suite and RoboDesk showcase the superior performance of our approach over baselines.
翻译:从高维、嘈杂且非平稳的环境中提取稳定且紧凑的表示,对于高效强化学习至关重要。此类环境中并存着不同类别的信息——如何有效提取并解耦这些信息仍是一项具有挑战性的问题。本文提出IFactor框架,该框架基于潜在状态变量与动作及回报的交互关系,将强化学习系统中的潜在状态变量建模为四类不同类别,从而捕获系统中各类信息。我们的分析建立了这些潜在变量的分块可识别性,这不仅提供了稳定且紧凑的表示,还揭示出所有与回报相关的因子对策略学习均具有重要意义。我们进一步提出一种实用方法,通过学习具有可识别分块的世界模型,确保去除冗余信息的同时保留策略优化的最小充分信息。在合成世界中的实验表明,我们的方法能准确识别真实潜在变量,验证了理论发现。此外,在DeepMind控制套件及RoboDesk变体环境中的实验显示,我们的方法相较于基线方法具有更优性能。