For a control problem with multiple conflicting objectives, there exists a set of Pareto-optimal policies called the Pareto set instead of a single optimal policy. When a multi-objective control problem is continuous and complex, traditional multi-objective reinforcement learning (MORL) algorithms search for many Pareto-optimal deep policies to approximate the Pareto set, which is quite resource-consuming. In this paper, we propose a simple and resource-efficient MORL algorithm that learns a continuous representation of the Pareto set in a high-dimensional policy parameter space using a single hypernet. The learned hypernet can directly generate various well-trained policy networks for different user preferences. We compare our method with two state-of-the-art MORL algorithms on seven multi-objective continuous robot control problems. Experimental results show that our method achieves the best overall performance with the least training parameters. An interesting observation is that the Pareto set is well approximated by a curved line or surface in a high-dimensional parameter space. This observation will provide insight for researchers to design new MORL algorithms.
翻译:对于具有多个冲突目标(目标)的控制问题,存在一组称为帕累托集的帕累托最优策略,而非单一最优策略。当多目标控制问题是连续且复杂时,传统的多目标强化学习(MORL)算法通过搜索大量帕累托最优深度策略来逼近帕累托集,这相当耗费资源。在本文中,我们提出了一种简单且资源高效的MORL算法,该算法使用单个超网络在高维策略参数空间中学习帕累托集的连续表示。学习到的超网络可以直接为不同的用户偏好生成各种训练良好的策略网络。我们在七个多目标连续机器人控制问题上,将我们的方法与两种最先进的MORL算法进行了比较。实验结果表明,我们的方法以最少的训练参数实现了最佳的整体性能。一个有趣的观察是,帕累托集在高维参数空间中很好地被一条曲线或曲面所逼近。这一观察将为研究人员设计新的MORL算法提供见解。