Reinforcement learning suffers from limitations in real practices primarily due to the numbers of required interactions with virtual environments. It results in a challenging problem that we are implausible to obtain an optimal strategy only with a few attempts for many learning method. Hereby, we design an improved reinforcement learning method based on model predictive control that models the environment through a data-driven approach. Based on learned environmental model, it performs multi-step prediction to estimate the value function and optimize the policy. The method demonstrates higher learning efficiency, faster convergent speed of strategies tending to the optimal value, and fewer sample capacity space required by experience replay buffers. Experimental results, both in classic databases and in a dynamic obstacle avoidance scenario for unmanned aerial vehicle, validate the proposed approaches.
翻译:强化学习在实际应用中受到限制,主要原因是需要与虚拟环境进行大量交互。这导致了一个具有挑战性的问题:对于许多学习方法而言,仅通过少量尝试难以获得最优策略。为此,我们设计了一种基于模型预测控制的改进强化学习方法,该方法通过数据驱动方式对环境进行建模。基于学到的环境模型,它执行多步预测以估计价值函数并优化策略。该方法展现出更高的学习效率、更快的策略收敛至最优值的速度,以及经验回放缓冲区所需更少的样本容量空间。在经典数据集和无人机动态避障场景中的实验结果均验证了所提方法的有效性。