Advancements in reinforcement learning (RL) have demonstrated superhuman performance in complex tasks such as Starcraft, Go, Chess etc. However, knowledge transfer from Artificial "Experts" to humans remain a significant challenge. A promising avenue for such transfer would be the use of curricula. Recent methods in curricula generation focuses on training RL agents efficiently, yet such methods rely on surrogate measures to track student progress, and are not suited for training robots in the real world (or more ambitiously humans). In this paper, we introduce a method named Parameterized Environment Response Model (PERM) that shows promising results in training RL agents in parameterized environments. Inspired by Item Response Theory, PERM seeks to model difficulty of environments and ability of RL agents directly. Given that RL agents and humans are trained more efficiently under the "zone of proximal development", our method generates a curriculum by matching the difficulty of an environment to the current ability of the student. In addition, PERM can be trained offline and does not employ non-stationary measures of student ability, making it suitable for transfer between students. We demonstrate PERM's ability to represent the environment parameter space, and training with RL agents with PERM produces a strong performance in deterministic environments. Lastly, we show that our method is transferable between students, without any sacrifice in training quality.
翻译:强化学习(RL)的最新进展已在星际争霸、围棋、国际象棋等复杂任务中展现出超人性能。然而,将人工智能“专家”的知识迁移至人类仍是一项重大挑战。实现此类迁移的一个有前景的途径是利用课程学习。近年来,课程生成方法主要关注高效训练RL智能体,但这些方法依赖替代指标来追踪学生进度,并不适用于真实世界机器人(或更宏大的目标——人类)的训练。本文提出一种名为参数化环境响应模型(PERM)的方法,其在参数化环境中训练RL智能体方面展现出令人鼓舞的结果。受项目反应理论启发,PERM旨在直接建模环境难度与智能体能力。鉴于RL智能体与人类在“最近发展区”内训练效率更高,我们的方法通过匹配环境难度与当前学生能力来生成课程。此外,PERM可离线训练,且不采用非平稳的学生能力度量,因此适用于不同学生间的迁移。我们证明了PERM对环境参数空间的表征能力,并在确定性环境中使用PERM训练RL智能体取得了优异性能。最后,我们展示该方法可在学生间实现可迁移性,且不牺牲训练质量。