This study evaluates the application of a discrete action space reinforcement learning method (Q-learning) to the continuous control problem of robot inverted pendulum balancing. To speed up the learning process and to overcome technical difficulties related to the direct learning on the real robotic system, the learning phase is performed in simulation environment. A mathematical model of the system dynamics is implemented, deduced by curve fitting on data acquired from the real system. The proposed approach demonstrated feasible, featuring its application on a real world robot that learned to balance an inverted pendulum. This study also reinforces and demonstrates the importance of an accurate representation of the physical world in simulation to achieve a more efficient implementation of reinforcement learning algorithms in real world, even when using a discrete action space algorithm to control a continuous action.
翻译:本研究评估了离散动作空间强化学习方法(Q-learning)在机器人倒立摆平衡连续控制问题中的应用。为加速学习过程并克服在真实机器人系统上直接学习的技术困难,学习阶段在仿真环境中进行。通过曲线拟合法从真实系统采集数据,推导并实现了系统动力学的数学模型。所提方法被证明切实可行,并在真实机器人上成功实现了倒立摆平衡学习。本研究同时强调并证明了:即使采用离散动作空间算法控制连续动作,物理世界的精确仿真表征对于强化学习算法在真实世界中的高效实现仍具有关键意义。