Model merging combines knowledge from separately fine-tuned models, yet the factors driving its success remain poorly understood. While recent work treats mergeability as an intrinsic property of the models, we show with an architecture-agnostic framework that it fundamentally depends on both the merging method and the partner tasks. Using L1-regularized linear optimization over a set of interpretable pairwise metrics (e.g., gradient $L_2$ distance), we uncover properties correlating with post-merge normalized accuracy across five merging methods. We find architecture- and method-specific variation in success drivers (64.0% average top-5 metric overlap; 79.3% sign agreement), with certain methods, notably TIES, exhibiting distinct ``fingerprints'' that diverge from the broader consensus. Crucially, however, \textit{gradient alignment} metrics consistently emerge as the most fundamental signals of compatibility. These findings provide a diagnostic foundation for understanding mergeability and motivate future merge-aware fine-tuning strategies.


翻译:模型融合结合了分别微调模型的知识,但决定其成功的关键因素尚不明确。尽管近期研究将融合性视为模型的内在属性,我们通过一个与架构无关的框架证明,融合性从根本上取决于融合方法和伙伴任务。基于一组可解释的成对度量指标(例如梯度$L_2$距离)的L1正则化线性优化,我们揭示了五种融合方法中与融合后归一化准确率相关的属性。我们发现成功驱动因素存在架构特异性和方法特异性(平均前5个度量重叠率64.0%,符号一致性79.3%),其中某些方法(特别是TIES)表现出偏离主流共识的独特“指纹”。然而关键的是,\textit{梯度对齐}度量始终是兼容性最基础的信号。这些发现为理解融合性提供了诊断基础,并启发了未来融合感知的微调策略。

0
下载
关闭预览

相关内容

融合知识图谱的预训练模型研究综述
专知会员服务
49+阅读 · 2024年3月31日
《深度模型融合》综述
专知会员服务
76+阅读 · 2023年9月28日
基于深度学习的图像融合方法综述
专知会员服务
58+阅读 · 2023年1月25日
基于深度学习的数据融合方法研究综述
专知会员服务
148+阅读 · 2020年12月10日
专知会员服务
224+阅读 · 2020年8月1日
基于深度学习的数据融合方法研究综述
专知
37+阅读 · 2020年12月10日
深度学习模型可解释性的研究进展
专知
26+阅读 · 2020年8月1日
专家报告|深度学习+图像多模态融合
中国图象图形学报
12+阅读 · 2019年10月23日
如何理解模型的过拟合与欠拟合,以及如何解决?
七月在线实验室
12+阅读 · 2019年4月23日
用模型不确定性理解模型
论智
11+阅读 · 2018年9月5日
【学界】机器学习模型的“可解释性”到底有多重要?
GAN生成式对抗网络
12+阅读 · 2018年3月3日
机器学习模型的“可解释性”到底有多重要?
中国科学院自动化研究所
20+阅读 · 2018年3月1日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
14+阅读 · 2023年9月27日
VIP会员
相关主题
最新内容
失去控制的指挥:人工智能时代的任务式指挥
专知会员服务
8+阅读 · 9月11日
美国的新国家安全科技战略思考
专知会员服务
4+阅读 · 9月11日
综述 | 面向大模型智能体的图结构个性化记忆
专知会员服务
7+阅读 · 9月10日
相关资讯
基于深度学习的数据融合方法研究综述
专知
37+阅读 · 2020年12月10日
深度学习模型可解释性的研究进展
专知
26+阅读 · 2020年8月1日
专家报告|深度学习+图像多模态融合
中国图象图形学报
12+阅读 · 2019年10月23日
如何理解模型的过拟合与欠拟合,以及如何解决?
七月在线实验室
12+阅读 · 2019年4月23日
用模型不确定性理解模型
论智
11+阅读 · 2018年9月5日
【学界】机器学习模型的“可解释性”到底有多重要?
GAN生成式对抗网络
12+阅读 · 2018年3月3日
机器学习模型的“可解释性”到底有多重要?
中国科学院自动化研究所
20+阅读 · 2018年3月1日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员