We study matrix estimation problems arising in reinforcement learning (RL) with low-rank structure. In low-rank bandits, the matrix to be recovered specifies the expected arm rewards, and for low-rank Markov Decision Processes (MDPs), it may for example characterize the transition kernel of the MDP. In both cases, each entry of the matrix carries important information, and we seek estimation methods with low entry-wise error. Importantly, these methods further need to accommodate for inherent correlations in the available data (e.g. for MDPs, the data consists of system trajectories). We investigate the performance of simple spectral-based matrix estimation approaches: we show that they efficiently recover the singular subspaces of the matrix and exhibit nearly-minimal entry-wise error. These new results on low-rank matrix estimation make it possible to devise reinforcement learning algorithms that fully exploit the underlying low-rank structure. We provide two examples of such algorithms: a regret minimization algorithm for low-rank bandit problems, and a best policy identification algorithm for reward-free RL in low-rank MDPs. Both algorithms yield state-of-the-art performance guarantees.
翻译:我们研究强化学习中具有低秩结构的矩阵估计问题。在低秩赌博机问题中,待恢复矩阵指定了期望的臂奖励值;而在低秩马尔可夫决策过程中,该矩阵可表征转移核等关键信息。在这两种情形下,矩阵的每个分量都承载重要信息,因此我们需要建立具有低分量级误差的估计方法。更重要的是,这些方法还需适应数据中固有的相关性(例如对马尔可夫决策过程而言,数据包含系统轨迹)。本文考察了简单谱方法在矩阵估计中的表现:我们证明该方法能高效恢复矩阵的奇异子空间,并实现近乎最优的分量级误差。这些关于低秩矩阵估计的新成果,使得开发充分利用底层低秩结构的强化学习算法成为可能。我们提供了两类此类算法:针对低秩赌博机问题的遗憾最小化算法,以及面向低秩马尔可夫决策过程的无奖励学习中最优策略识别算法。两种算法均实现了当前最优的性能保证。