While much progress has been made in understanding the minimax sample complexity of reinforcement learning (RL) -- the complexity of learning on the "worst-case" instance -- such measures of complexity often do not capture the true difficulty of learning. In practice, on an "easy" instance, we might hope to achieve a complexity far better than that achievable on the worst-case instance. In this work we seek to understand the "instance-dependent" complexity of learning near-optimal policies (PAC RL) in the setting of RL with linear function approximation. We propose an algorithm, \textsc{Pedel}, which achieves a fine-grained instance-dependent measure of complexity, the first of its kind in the RL with function approximation setting, thereby capturing the difficulty of learning on each particular problem instance. Through an explicit example, we show that \textsc{Pedel} yields provable gains over low-regret, minimax-optimal algorithms and that such algorithms are unable to hit the instance-optimal rate. Our approach relies on a novel online experiment design-based procedure which focuses the exploration budget on the "directions" most relevant to learning a near-optimal policy, and may be of independent interest.
翻译:尽管在理解强化学习(RL)的极小极大样本复杂度——即学习"最坏情况"实例的复杂度——方面取得了诸多进展,但这种复杂度度量通常无法捕捉学习的真实难度。在实践中,对于"简单"实例,我们可能希望获得远优于最坏情况实例的复杂度。本研究旨在理解线性函数逼近强化学习设置下学习近最优策略(PAC RL)的"实例依赖"复杂度。我们提出算法\textsc{Pedel},该算法实现了细粒度的实例依赖复杂度度量——这是在线性函数逼近RL设置中的首次此类成果,从而能够捕捉每个特定问题实例的学习难度。通过一个显式示例,我们证明\textsc{Pedel}相比低遗憾、极小极大最优算法具有可证明的性能提升,且此类算法无法达到实例最优速率。我们的方法依赖于一种新颖的基于在线实验设计的程序,该程序将探索预算集中于与学习近最优策略最相关的"方向"上,该技术本身可能具有独立的研究价值。