Inspired by Regularized Lottery Ticket Hypothesis (RLTH), which states that competitive smooth (non-binary) subnetworks exist within a dense network in continual learning tasks, we investigate two proposed architecture-based continual learning methods which sequentially learn and select adaptive binary- (WSN) and non-binary Soft-Subnetworks (SoftNet) for each task. WSN and SoftNet jointly learn the regularized model weights and task-adaptive non-binary masks of subnetworks associated with each task whilst attempting to select a small set of weights to be activated (winning ticket) by reusing weights of the prior subnetworks. Our proposed WSN and SoftNet are inherently immune to catastrophic forgetting as each selected subnetwork model does not infringe upon other subnetworks in Task Incremental Learning (TIL). In TIL, binary masks spawned per winning ticket are encoded into one N-bit binary digit mask, then compressed using Huffman coding for a sub-linear increase in network capacity to the number of tasks. Surprisingly, in the inference step, SoftNet generated by injecting small noises to the backgrounds of acquired WSN (holding the foregrounds of WSN) provides excellent forward transfer power for future tasks in TIL. SoftNet shows its effectiveness over WSN in regularizing parameters to tackle the overfitting, to a few examples in Few-shot Class Incremental Learning (FSCIL).
翻译:受正则化彩票假设(RLTH)的启发——该假设指出在持续学习任务中,稠密网络内存在具有竞争力的平滑(非二值)子网络——我们研究了两种基于架构的持续学习方法,它们依次学习并为每个任务选择自适应的二值(WSN)和非二值软子网络(SoftNet)。WSN和SoftNet联合学习正则化模型权重以及与每个任务相关的任务自适应非二值子网络掩码,同时尝试通过重用先前子网络的权重来选择一小组待激活的权重(胜出票)。我们提出的WSN和SoftNet本质上能够抵抗灾难性遗忘,因为在任务增量学习(TIL)中,每个选定的子网络模型不会侵犯其他子网络。在TIL中,每个胜出票产生的二值掩码被编码为单个N位二值数字掩码,然后使用霍夫曼编码进行压缩,使得网络容量随任务数量呈亚线性增长。令人惊讶的是,在推理阶段,通过向所获取WSN的背景(保留WSN的前景)注入小噪声生成的SoftNet,为TIL中的未来任务提供了出色的前向迁移能力。在小样本类别增量学习(FSCIL)中,SoftNet在针对少数样本的正则化参数防止过拟合方面,展现出优于WSN的有效性。