In modern machine learning, models can often fit training data in numerous ways, some of which perform well on unseen (test) data, while others do not. Remarkably, in such cases gradient descent frequently exhibits an implicit bias that leads to excellent performance on unseen data. This implicit bias was extensively studied in supervised learning, but is far less understood in optimal control (reinforcement learning). There, learning a controller applied to a system via gradient descent is known as policy gradient, and a question of prime importance is the extent to which a learned controller extrapolates to unseen initial states. This paper theoretically studies the implicit bias of policy gradient in terms of extrapolation to unseen initial states. Focusing on the fundamental Linear Quadratic Regulator (LQR) problem, we establish that the extent of extrapolation depends on the degree of exploration induced by the system when commencing from initial states included in training. Experiments corroborate our theory, and demonstrate its conclusions on problems beyond LQR, where systems are non-linear and controllers are neural networks. We hypothesize that real-world optimal control may be greatly improved by developing methods for informed selection of initial states to train on.
翻译:在现代机器学习中,模型通常能以多种方式拟合训练数据,其中一些方式在未见(测试)数据上表现良好,而另一些则不然。值得注意的是,在这种情况下,梯度下降经常表现出一种隐式偏差,从而在未见数据上获得优异性能。这种隐式偏差在监督学习中得到了广泛研究,但在最优控制(强化学习)中却远未得到充分理解。在该领域,通过梯度下降学习应用于系统的控制器被称为策略梯度,一个至关重要的问题是学习到的控制器在多大程度上能外推至未见初始状态。本文从向未见初始状态外推的角度,理论上研究了策略梯度的隐式偏差。聚焦于基本的线性二次调节器(LQR)问题,我们证明了外推的程度取决于系统从训练包含的初始状态开始时引发的探索程度。实验证实了我们的理论,并在超越LQR的问题(系统是非线性的且控制器是神经网络)上展示了其结论。我们假设,通过开发用于知情选择训练初始状态的方法,现实世界的最优控制可能会得到极大改进。