Task arithmetic has recently emerged as a cost-effective and scalable approach to edit pre-trained models directly in weight space: By adding the fine-tuned weights of different tasks, the model's performance can be improved on these tasks, while negating them leads to task forgetting. Yet, our understanding of the effectiveness of task arithmetic and its underlying principles remains limited. We present a comprehensive study of task arithmetic in vision-language models and show that weight disentanglement is the crucial factor that makes it effective. This property arises during pre-training and manifests when distinct directions in weight space govern separate, localized regions in function space associated with the tasks. Notably, we show that fine-tuning models in their tangent space by linearizing them amplifies weight disentanglement. This leads to substantial performance improvements across multiple task arithmetic benchmarks and diverse models. Building on these findings, we provide theoretical and empirical analyses of the neural tangent kernel (NTK) of these models and establish a compelling link between task arithmetic and the spatial localization of the NTK eigenfunctions. Overall, our work uncovers novel insights into the fundamental mechanisms of task arithmetic and offers a more reliable and effective approach to edit pre-trained models through the NTK linearization.
翻译:任务算术近年来作为一种经济且可扩展的方法,直接在权重空间中对预训练模型进行编辑:通过叠加不同任务的微调权重,可提升模型在这些任务上的性能,而对这些权重取反则会导致任务遗忘。然而,我们对任务算术有效性及其基本原理的理解仍然有限。我们针对视觉-语言模型中的任务算术开展了全面研究,并发现权重解缠是使其有效的关键因素。这一性质在预训练阶段形成,其表现为权重空间中的不同方向控制着函数空间中与任务相关的独立局部区域。值得注意的是,我们证明通过线性化模型在切空间中进行微调,能够增强权重解缠。这显著提升了任务算术在多个基准测试及不同模型上的性能。基于这些发现,我们对这些模型的神经正切核(NTK)进行了理论和实证分析,并建立了任务算术与NTK本征函数空间局部化之间的紧密联系。总体而言,我们的研究揭示了任务算术基本机制的新见解,并通过NTK线性化为预训练模型提供了一种更可靠、更有效的编辑方法。