Automatic Speech Recognition (ASR) systems have attained unprecedented performance with large speech models pre-trained based on self-supervised speech representation learning. However, these pre-trained speech models suffer from representational bias as they tend to better represent those prominent accents (i.e., native (L1) English accent) in the pre-training speech corpus than less represented accents, resulting in a deteriorated performance for non-native (L2) English accents. Although there have been some approaches to mitigate this issue, all of these methods require updating the pre-trained model weights. In this paper, we propose Information Theoretic Adversarial Prompt Tuning (INTapt), which introduces prompts concatenated to the original input that can re-modulate the attention of the pre-trained model such that the corresponding input resembles a native (L1) English speech without updating the backbone weights. INTapt is trained simultaneously in the following two manners: (1) adversarial training to reduce accent feature dependence between the original input and the prompt-concatenated input and (2) training to minimize CTC loss for improving ASR performance to a prompt-concatenated input. Experimental results show that INTapt improves the performance of L2 English and increases feature similarity between L2 and L1 accents.
翻译:自动语音识别(ASR)系统通过基于自监督语音表示学习的预训练大型语音模型已取得了前所未有的性能。然而,这些预训练语音模型存在表示偏差,因为它们倾向于更好地表示预训练语音语料库中那些突出的口音(即母语(L1)英语口音),而对代表性不足的口音(即非母语(L2)英语口音)表现不佳,导致其性能下降。尽管已有一些方法来缓解这一问题,但这些方法均需要更新预训练模型权重。本文提出信息论对抗提示调优(INTapt),该方法引入拼接在原始输入上的提示,能够重新调制预训练模型的注意力,使相应输入在不更新骨干网络权重的情况下类似于母语(L1)英语语音。INTapt同时通过以下两种方式进行训练:(1)对抗训练,以降低原始输入与提示拼接输入之间的口音特征依赖性;(2)训练以最小化CTC损失,从而提升提示拼接输入的ASR性能。实验结果表明,INTapt改善了L2英语的性能,并增加了L2与L1口音之间的特征相似性。