Motivated by the challenge of achieving rapid learning in physical environments, this paper presents the development and training of a robotic system designed to navigate and solve a labyrinth game using model-based reinforcement learning techniques. The method involves extracting low-dimensional observations from camera images, along with a cropped and rectified image patch centered on the current position within the labyrinth, providing valuable information about the labyrinth layout. The learning of a control policy is performed purely on the physical system using model-based reinforcement learning, where the progress along the labyrinth's path serves as a reward signal. Additionally, we exploit the system's inherent symmetries to augment the training data. Consequently, our approach learns to successfully solve a popular real-world labyrinth game in record time, with only 5 hours of real-world training data.
翻译:受物理环境中快速学习挑战的驱动,本文介绍了一种基于模型强化学习技术的机器人系统的开发与训练,该系统旨在导航并解决迷宫游戏。该方法从相机图像中提取低维观测值,并结合以迷宫内当前位置为中心的裁剪和校正图像块,提供有关迷宫布局的有价值信息。控制策略的学习完全在物理系统上通过模型强化学习进行,其中沿迷宫路径的进展作为奖励信号。此外,我们利用系统固有的对称性来增强训练数据。因此,我们的方法在创纪录的时间内成功学会了解决一个流行的真实迷宫游戏,仅使用5小时的真实训练数据。