Deep neural networks are susceptible to catastrophic forgetting when trained on sequential tasks. Various continual learning (CL) methods often rely on exemplar buffers or/and network expansion for balancing model stability and plasticity, which, however, compromises their practical value due to privacy and memory concerns. Instead, this paper considers a strict yet realistic setting, where the training data from previous tasks is unavailable and the model size remains relatively constant during sequential training. To achieve such desiderata, we propose a conceptually simple yet effective method that attributes forgetting to layer-wise parameter overwriting and the resulting decision boundary distortion. This is achieved by the synergy between two key components: HSIC-Bottleneck Orthogonalization (HBO) implements non-overwritten parameter updates mediated by Hilbert-Schmidt independence criterion in an orthogonal space and EquiAngular Embedding (EAE) enhances decision boundary adaptation between old and new tasks with predefined basis vectors. Extensive experiments demonstrate that our method achieves competitive accuracy performance, even with absolute superiority of zero exemplar buffer and 1.02x the base model.
翻译:深度神经网络在序列任务训练时易遭受灾难性遗忘。现有持续学习方法常依赖示例缓冲区或/和网络扩展来平衡模型稳定性与可塑性,但由于隐私和内存问题而损害其实用价值。为此,本文考虑一种严格且现实的情景:先前任务的训练数据不可获取,且模型规模在序列训练过程中保持相对恒定。为实现此类理想目标,我们提出一种概念简洁而有效的方法,将遗忘归因于逐层参数覆写及其导致的决策边界畸变。该方法通过两个关键组件的协同作用实现:HSIC瓶颈正交化(HBO)在正交空间中通过希尔伯特-施密特独立性准则实现非覆写参数更新,等角嵌入(EAE)利用预定义基向量增强新旧任务间的决策边界自适应。大量实验表明,本方法在保持零示例缓冲区和基模型1.02倍参数量的绝对优势下,仍能达到具有竞争力的准确率性能。