Deep reinforcement learning policies, despite their outstanding efficiency in simulated visual control tasks, have shown disappointing ability to generalize across disturbances in the input training images. Changes in image statistics or distracting background elements are pitfalls that prevent generalization and real-world applicability of such control policies. We elaborate on the intuition that a good visual policy should be able to identify which pixels are important for its decision, and preserve this identification of important sources of information across images. This implies that training of a policy with small generalization gap should focus on such important pixels and ignore the others. This leads to the introduction of saliency-guided Q-networks (SGQN), a generic method for visual reinforcement learning, that is compatible with any value function learning method. SGQN vastly improves the generalization capability of Soft Actor-Critic agents and outperforms existing stateof-the-art methods on the Deepmind Control Generalization benchmark, setting a new reference in terms of training efficiency, generalization gap, and policy interpretability.
翻译:深度强化学习策略虽然在模拟视觉控制任务中表现出色,但在面对输入训练图像中的扰动时,其泛化能力却令人失望。图像统计特征的变化或分散注意力的背景元素是阻碍此类控制策略泛化及实际应用的关键因素。我们基于一个直觉性认识展开研究:优秀的视觉策略应能识别哪些像素对决策至关重要,并在不同图像中保持对这种重要信息源的识别能力。这意味着,具有较小泛化差距的策略训练应聚焦于这些关键像素而忽略其他信息。基于此,我们提出显著性引导的Q网络(SGQN)——一种通用的视觉强化学习方法,可与任意价值函数学习方法兼容。SGQN显著提升了软演员-评论家(Soft Actor-Critic)智能体的泛化能力,在Deepmind控制泛化基准测试中超越了现有最先进方法,在训练效率、泛化差距和策略可解释性方面树立了新的标杆。