We present AIRS: Automatic Intrinsic Reward Shaping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task return in real-time, providing reliable exploration incentives and alleviating the biased objective problem. Moreover, we develop an intrinsic reward toolkit to provide efficient and reliable implementations of diverse intrinsic reward approaches. We test AIRS on various tasks of MiniGrid, Procgen, and DeepMind Control Suite. Extensive simulation demonstrates that AIRS can outperform the benchmarking schemes and achieve superior performance with simple architecture.
翻译:我们提出AIRS:一种自动内在奖励塑形方法,能够智能且自适应地提供高质量的内在奖励,以增强强化学习中的探索能力。具体而言,AIRS基于实时估计的任务回报,从预定义集合中选择塑形函数,从而提供可靠的探索激励并缓解目标偏差问题。此外,我们开发了一个内在奖励工具包,以高效且可靠地实现多种内在奖励方法。我们在MiniGrid、Procgen和DeepMind Control Suite的各类任务上测试了AIRS。大量仿真结果表明,AIRS能够优于基线方案,并以简洁架构实现卓越性能。