End-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) aims to improve the performance on the unlabeled target domain by transferring knowledge from the source to the target domain. To improve transferability, existing UDA approaches mainly focus on matching the distributions of the source and target domains globally and/or locally, while ignoring the model discriminability. In this paper, we propose a novel UDA approach for ASR via inter-domain MAtching and intra-domain DIscrimination (MADI), which improves the model transferability by fine-grained inter-domain matching and discriminability by intra-domain contrastive discrimination simultaneously. Evaluations on the Libri-Adapt dataset demonstrate the effectiveness of our approach. MADI reduces the relative word error rate (WER) on cross-device and cross-environment ASR by 17.7% and 22.8%, respectively.
翻译:端到端自动语音识别(ASR)在应用于新领域时,常因领域偏移导致性能下降。无监督域适应(UDA)旨在通过将知识从源领域迁移至目标领域,提升在未标注目标领域上的性能。现有UDA方法主要聚焦于全局和/或局部层面匹配源域与目标域的分布,以提升可迁移性,却忽略了模型的判别能力。本文提出了一种面向ASR的新型UDA方法——域间匹配与域内判别(MADI),该方法通过细粒度的域间匹配提升模型可迁移性,同时通过域内对比判别增强判别能力。在Libri-Adapt数据集上的评估验证了本方法的有效性。MADI在跨设备与跨环境ASR任务上分别降低了17.7%和22.8%的相对词错误率(WER)。