Modern representation learning methods may fail to adapt quickly under non-stationarity since they suffer from the problem of catastrophic forgetting and decaying plasticity. Such problems prevent learners from fast adaptation to changes since they result in increasing numbers of saturated features and forgetting useful features when presented with new experiences. Hence, these methods are rendered ineffective for continual learning. This paper proposes Utility-based Perturbed Gradient Descent (UPGD), an online representation-learning algorithm well-suited for continual learning agents with no knowledge about task boundaries. UPGD protects useful weights or features from forgetting and perturbs less useful ones based on their utilities. Our empirical results show that UPGD alleviates catastrophic forgetting and decaying plasticity, enabling modern representation learning methods to work in the continual learning setting.
翻译:现代表示学习方法在非平稳环境下可能难以快速适应,因为它们面临灾难性遗忘和可塑性衰退的问题。这些问题导致学习者在面对新经验时产生大量饱和特征并遗忘有用特征,从而阻碍其快速适应变化。因此,这些方法在持续学习中效果不佳。本文提出基于效用的扰动梯度下降(UPGD),这是一种在线表示学习算法,特别适用于未知任务边界的持续学习智能体。UPGD根据特征的效用保护有用权重或特征免于遗忘,同时扰动效用较低的特征。我们的实验结果表明,UPGD能够缓解灾难性遗忘和可塑性衰退,使现代表示学习方法能够在持续学习场景中有效工作。