Foundation models encompass an extensive knowledge base and offer remarkable transferability. However, this knowledge becomes outdated or insufficient over time. The challenge lies in continuously updating foundation models to accommodate novel information while retaining their original capabilities. Leveraging the fact that foundation models have initial knowledge on various tasks and domains, we propose a novel approach that, instead of updating all parameters equally, localizes the updates to a sparse set of parameters relevant to the task being learned. We strike a balance between efficiency and new task performance, while maintaining the transferability and generalizability of foundation models. We extensively evaluate our method on foundational vision-language models with a diverse spectrum of continual learning tasks. Our method achieves improvements on the accuracy of the newly learned tasks up to 7% while preserving the pretraining knowledge with a negligible decrease of 0.9% on a representative control set accuracy.
翻译:基础模型蕴含广泛的知识库并具备显著的迁移能力。然而,随着时间的推移,这些知识可能变得过时或不足。当前挑战在于持续更新基础模型以吸收新信息,同时保留其原有能力。利用基础模型在各类任务和领域上具有先验知识这一特性,我们提出一种新方法:并非对所有参数进行等量更新,而是将更新过程定位到与当前学习任务相关的稀疏参数集合上。该方法在效率与新任务性能之间取得平衡,同时保持基础模型的迁移性和泛化能力。我们针对基础视觉语言模型开展了一系列持续学习任务的全面评估。实验表明,本方法在新任务准确率上最高提升7%,而预训练知识的保留仅损失0.9%(以具有代表性的控制集准确率衡量),降幅可忽略不计。