On-policy algorithms are supposed to be stable, however, sample-intensive yet. Off-policy algorithms utilizing past experiences are deemed to be sample-efficient, nevertheless, unstable in general. Can we design an algorithm that can employ the off-policy data, while exploit the stable learning by sailing along the course of the on-policy walkway? In this paper, we present an actor-critic learning framework that borrows the distributional perspective of interest to evaluate, and cross-breeds two sources of the data for policy improvement, which enables fast learning and can be applied to a wide class of algorithms. In its backbone, the variance reduction mechanisms, such as unified advantage estimator (UAE), that extends generalized advantage estimator (GAE) to be applicable on any state-dependent baseline, and a learned baseline, that is competent to stabilize the policy gradient, are firstly put forward to not merely be a bridge to the action-value function but also distill the advantageous learning signal. Lastly, it is empirically shown that our method improves sample efficiency and interpolates different levels well. Being of an organic whole, its mixture places more inspiration to the algorithm design.
翻译:在线策略算法理论上具有稳定性,但采样效率低下;离线策略算法利用历史经验具有样本高效性,但通常不稳定。能否设计一种算法,既能利用离线策略数据,又能通过沿在线策略路径航行来利用稳定学习?本文提出了一种演员-评论家学习框架,该框架借鉴了感兴趣的分布视角进行评估,并通过交叉融合两种数据源进行策略改进,从而实现快速学习,并可应用于广泛算法类别。在其核心部分,首先提出了方差缩减机制,如统一优势估计器(UAE),该机制将广义优势估计器(GAE)扩展到适用于任何状态相关基线,以及一种能够稳定策略梯度的学习基线,这些机制不仅充当连接动作值函数的桥梁,还提取了有利的学习信号。最后,实证表明我们的方法提高了样本效率,并很好地插值了不同水平。作为一个有机整体,其融合为算法设计提供了更多灵感。