There has recently been an increased interest in reinforcement learning for nonlinear control problems. However standard reinforcement learning algorithms can often struggle even on seemingly simple set-point control problems. This paper argues that three ideas can improve reinforcement learning methods even for highly nonlinear set-point control problems: 1) Make use of a prior feedback controller to aid amplitude exploration. 2) Use integrated errors. 3) Train on model ensembles. Together these ideas lead to more efficient training, and a trained set-point controller that is more robust to modelling errors and thus can be directly deployed to real-world nonlinear systems. The claim is supported by experiments with a real-world nonlinear cascaded tank process and a simulated strongly nonlinear pH-control system.
翻译:近年来,强化学习在非线性控制问题中的应用备受关注。然而,即使面对看似简单的设定点控制问题,标准强化学习算法也常常难以奏效。本文论证了三种能够提升强化学习方法性能的理念,即便针对高度非线性的设定点控制问题也同样有效:1)利用先验反馈控制器辅助幅度探索;2)采用积分误差;3)基于模型集成进行训练。这些理念的结合能够实现更高效的训练,并使得训练出的设定点控制器对模型误差具有更强的鲁棒性,从而可直接部署于实际非线性系统。这一论断得到了真实非线性级联水箱过程实验以及模拟强非线性pH控制系统实验的支持。