Recent advances in visual reinforcement learning (RL) have led to impressive success in handling complex tasks. However, these methods have demonstrated limited generalization capability to visual disturbances, which poses a significant challenge for their real-world application and adaptability. Though normalization techniques have demonstrated huge success in supervised and unsupervised learning, their applications in visual RL are still scarce. In this paper, we explore the potential benefits of integrating normalization into visual RL methods with respect to generalization performance. We find that, perhaps surprisingly, incorporating suitable normalization techniques is sufficient to enhance the generalization capabilities, without any additional special design. We utilize the combination of two normalization techniques, CrossNorm and SelfNorm, for generalizable visual RL. Extensive experiments are conducted on DMControl Generalization Benchmark and CARLA to validate the effectiveness of our method. We show that our method significantly improves generalization capability while only marginally affecting sample efficiency. In particular, when integrated with DrQ-v2, our method enhances the test performance of DrQ-v2 on CARLA across various scenarios, from 14% of the training performance to 97%.
翻译:近期视觉强化学习(RL)的进展在应对复杂任务方面取得了显著成功。然而,这些方法在视觉扰动下表现出有限的泛化能力,这对其实际应用和适应性构成了重大挑战。尽管归一化技术在监督学习和无监督学习中取得了巨大成功,但在视觉RL中的应用仍然很少。本文探讨了在视觉RL方法中融入归一化对泛化性能的潜在益处。我们发现,出乎意料的是,采用合适的归一化技术足以增强泛化能力,而无需任何额外的特殊设计。我们利用两种归一化技术CrossNorm和SelfNorm的组合来实现可泛化的视觉RL。在DMControl泛化基准测试和CARLA上进行了大量实验,以验证我们方法的有效性。结果表明,我们的方法显著提升了泛化能力,同时仅对样本效率产生微弱影响。特别地,当与DrQ-v2集成时,我们的方法将DrQ-v2在CARLA上跨多种场景的测试性能从训练性能的14%提升至97%。