Although Reinforcement Learning (RL) has shown to be capable of producing impressive results, its use is limited by the impact of its hyperparameters on performance. This often makes it difficult to achieve good results in practice. Automated RL (AutoRL) addresses this difficulty, yet little is known about the dynamics of the hyperparameter landscapes that hyperparameter optimization (HPO) methods traverse in search of optimal configurations. In view of existing AutoRL approaches dynamically adjusting hyperparameter configurations, we propose an approach to build and analyze these hyperparameter landscapes not just for one point in time but at multiple points in time throughout training. Addressing an important open question on the legitimacy of such dynamic AutoRL approaches, we provide thorough empirical evidence that the hyperparameter landscapes strongly vary over time across representative algorithms from RL literature (DQN and SAC) in different kinds of environments (Cartpole and Hopper). This supports the theory that hyperparameters should be dynamically adjusted during training and shows the potential for more insights on AutoRL problems that can be gained through landscape analyses.
翻译:尽管强化学习(RL)已展现出产生令人瞩目成果的能力,但其应用仍受超参数对性能影响的制约,这常导致实践中难以获得理想结果。自动强化学习(AutoRL)旨在解决这一难题,然而目前人们对超参数优化(HPO)方法在搜索最优配置时所遍历的超参数景观动态特性知之甚少。鉴于现有AutoRL方法会动态调整超参数配置,我们提出一种方法,不仅针对单一时间点,更在训练过程中的多个时间点构建和分析这些超参数景观。针对此类动态AutoRL方法合法性的重要开放性问题,我们提供了充分的实证证据,证明在RL文献中的代表性算法(DQN与SAC)及不同类型环境(Cartpole与Hopper)中,超参数景观随时间剧烈变化。这一发现为训练过程中应动态调整超参数的理论提供了有力支撑,并展示了通过景观分析可能为AutoRL问题带来更多深刻洞见的前景。