We study the task of learning state representations from potentially high-dimensional observations, with the goal of controlling an unknown partially observable system. We pursue a direct latent model learning approach, where a dynamic model in some latent state space is learned by predicting quantities directly related to planning (e.g., costs) without reconstructing the observations. In particular, we focus on an intuitive cost-driven state representation learning method for solving Linear Quadratic Gaussian (LQG) control, one of the most fundamental partially observable control problems. As our main results, we establish finite-sample guarantees of finding a near-optimal state representation function and a near-optimal controller using the directly learned latent model. To the best of our knowledge, despite various empirical successes, prior to this work it was unclear if such a cost-driven latent model learner enjoys finite-sample guarantees. Our work underscores the value of predicting multi-step costs, an idea that is key to our theory, and notably also an idea that is known to be empirically valuable for learning state representations.
翻译:我们研究从潜在高维观测中学习状态表征的任务,目标是对未知部分可观测系统进行控制。我们采用直接潜变量模型学习方法,即通过预测直接与规划相关的量(如代价)来学习潜状态空间中的动态模型,而无需重构观测值。具体而言,我们聚焦于一种直观的代价驱动状态表征学习方法,用于解决线性二次高斯(LQG)控制问题——这是最基础的部分可观测控制问题之一。作为主要结论,我们建立了通过直接学习潜变量模型获得近似最优状态表征函数和近似最优控制器的有限样本保证。据我们所知,尽管已有诸多实证成功案例,但在此工作之前,这种代价驱动的潜变量模型学习器是否享有有限样本保证尚不明确。我们的工作凸显了预测多步代价的价值——这一概念既是理论的关键,也在经验上被证明对学习状态表征具有重要价值。