While self-supervised speech representation learning (SSL) models serve a variety of downstream tasks, these models have been observed to overfit to the domain from which the unlabelled data originates. To alleviate this issue, we propose PADA (Pruning Assisted Domain Adaptation) and zero out redundant weights from models pre-trained on large amounts of out-of-domain (OOD) data. Intuitively, this helps to make space for the target-domain ASR finetuning. The redundant weights can be identified through various pruning strategies which have been discussed in detail as a part of this work. Specifically, we investigate the effect of the recently discovered Task-Agnostic and Task-Aware pruning on PADA and propose a new pruning paradigm based on the latter, which we call Cross-Domain Task-Aware Pruning (CD-TAW). CD-TAW obtains the initial pruning mask from a well fine-tuned OOD model, which makes it starkly different from the rest of the pruning strategies discussed in the paper. Our proposed CD-TAW methodology achieves up to 20.6% relative WER improvement over our baseline when fine-tuned on a 2-hour subset of Switchboard data without language model (LM) decoding. Furthermore, we conduct a detailed analysis to highlight the key design choices of our proposed method.
翻译:摘要:尽管自监督语音表征学习(SSL)模型可服务于多种下游任务,但此类模型已被观察到会过度拟合于无标签数据所来源的领域。为解决这一问题,我们提出PADA(剪枝辅助域自适应),将基于大量域外(OOD)数据预训练的模型中的冗余权重归零。直观而言,这有助于为目标领域的ASR微调腾出空间。这些冗余权重可通过多种剪枝策略识别,本文对此进行了详细讨论。具体而言,我们研究了近期提出的任务无关剪枝与任务感知剪枝对PADA的影响,并基于后者提出一种新的剪枝范式——跨域任务感知剪枝(CD-TAW)。CD-TAW从经过良好微调的OOD模型中获取初始剪枝掩码,使其与本文讨论的其他剪枝策略存在显著差异。我们提出的CD-TAW方法在不使用语言模型(LM)解码的情况下,基于Switchboard数据集的2小时子集进行微调,相对于基线实现了最高20.6%的相对词错误率(WER)改进。此外,我们通过详细分析重点阐述了所提方法的关键设计选择。