ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29\% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
翻译:基于自监督声学预训练和CTC微调的ASR系统在母语语音上表现优异,但对口音变异性仍较敏感。本文研究将监督对比学习(SupCon)作为CTC微调过程中轻量级、口音不变性的辅助目标。在不修改架构或使用显式口音监督的情况下,语句级对比损失可对编码器表示进行正则化。在L2-ARCTIC基准上的实验表明,多种预训练编码器均实现一致的词错误率降低,在未见口音评估中相对降幅达25%-29%。基于转录内余弦离散度的分析表明,SupCon可促进口音变异性下更紧凑且稳定的表示几何结构。总体而言,SupCon为提升口音鲁棒性提供了有效且模型无关的正则化策略。