超越锐度：一种面向高效持续学习的平坦性分解框架 (Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual Learning)

Continual Learning (CL) aims to enable models to sequentially learn multiple tasks without forgetting previous knowledge. Recent studies have shown that optimizing towards flatter loss minima can improve model generalization. However, existing sharpness-aware methods for CL suffer from two key limitations: (1) they treat sharpness regularization as a unified signal without distinguishing the contributions of its components. and (2) they introduce substantial computational overhead that impedes practical deployment. To address these challenges, we propose FLAD, a novel optimization framework that decomposes sharpness-aware perturbations into gradient-aligned and stochastic-noise components, and show that retaining only the noise component promotes generalization. We further introduce a lightweight scheduling scheme that enables FLAD to maintain significant performance gains even under constrained training time. FLAD can be seamlessly integrated into various CL paradigms and consistently outperforms standard and sharpness-aware optimizers in diverse experimental settings, demonstrating its effectiveness and practicality in CL.

翻译：持续学习（CL）旨在使模型能够顺序学习多个任务而不遗忘先前知识。近期研究表明，优化至更平坦的损失极小值可以提升模型泛化能力。然而，现有面向持续学习的锐度感知方法存在两个关键局限：（1）它们将锐度正则化视为统一信号，未区分其各组成部分的贡献；（2）它们引入了大量计算开销，阻碍了实际部署。为应对这些挑战，我们提出FLAD，一种新颖的优化框架，将锐度感知扰动分解为梯度对齐分量和随机噪声分量，并证明仅保留噪声分量即可促进泛化。我们进一步引入一种轻量级调度方案，使FLAD即使在受限训练时间下也能保持显著的性能提升。FLAD可无缝集成到多种持续学习范式中，并在多样化的实验设置中持续优于标准及锐度感知优化器，证明了其在持续学习中的有效性和实用性。

相关内容

持续学习

关注 25

持续学习(continuallearning,CL) 是模拟大脑学习的过程,按照一定的顺序对连续非独立同分布的 (independentlyandidenticallydistributed,IID)流数据进行学习,进而根据任务的执行结果对模型进行增量式更新．持续学习的意义在于高效地转化和利用已经学过的知识来完成新任务的学习,并且能够极大程度地降低遗忘带来的问题．连续学习研究对智能计算系统自适应地适应环境改变具有重要的意义

【NeurIPS2024】超越冗余：信息感知的无监督多重图结构学习

专知会员服务

28+阅读 · 2024年9月29日

【超越消息传递:图神经网络的物理启发范式】Beyond Message Passing: a Physics-Inspired Paradigm for Graph Neural Networks

专知会员服务

17+阅读 · 2022年5月10日

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

专知会员服务

18+阅读 · 2022年3月19日

【CMU-Yuejie Chi等干货书】满足低秩矩阵分解的非凸优化综述，69页pdf，Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview

专知会员服务

33+阅读 · 2022年3月4日