Training large foundation models using self-supervised objectives on unlabeled data, followed by fine-tuning on downstream tasks, has emerged as a standard procedure. Unfortunately, the efficacy of this approach is often constrained by both limited fine-tuning compute and scarcity in labeled downstream data. We introduce Multimodal Attention Merging (MAM), an attempt that facilitates direct knowledge transfer from attention matrices of models rooted in high resource modalities, text and images, to those in resource-constrained domains, speech and audio, employing a zero-shot paradigm. MAM reduces the relative Word Error Rate (WER) of an Automatic Speech Recognition (ASR) model by up to 6.70%, and relative classification error of an Audio Event Classification (AEC) model by 10.63%. In cases where some data/compute is available, we present Learnable-MAM, a data-driven approach to merging attention matrices, resulting in a further 2.90% relative reduction in WER for ASR and 18.42% relative reduction in AEC compared to fine-tuning.
翻译:利用自监督目标函数在无标注数据上训练大型基础模型,随后在下游任务上进行微调已成为标准流程。然而,该方法的效果常常受限于微调计算资源不足和下游标注数据稀缺。我们提出多模态注意力合并(MAM),该方法通过零样本范式,实现从高资源模态(文本与图像)模型注意力矩阵到资源受限领域(语音与音频)模型的直接知识迁移。MAM使自动语音识别(ASR)模型的相对词错误率(WER)降低6.70%,音频事件分类(AEC)模型的相对分类错误率降低10.63%。在具备部分数据/计算资源的情况下,我们提出可学习的注意力合并方法(Learnable-MAM),这是一种数据驱动的注意力矩阵合并策略,相较于微调方法,进一步使ASR模型的相对WER降低2.90%,AEC模型的相对分类错误率降低18.42%。