Automatic speech recognition (ASR) systems are known to be sensitive to the sociolinguistic variability of speech data, in which gender plays a crucial role. This can result in disparities in recognition accuracy between male and female speakers, primarily due to the under-representation of the latter group in the training data. While in the context of hybrid ASR models several solutions have been proposed, the gender bias issue has not been explicitly addressed in end-to-end neural architectures. To fill this gap, we propose a data augmentation technique that manipulates the fundamental frequency (f0) and formants. This technique reduces the data unbalance among genders by simulating voices of the under-represented female speakers and increases the variability within each gender group. Experiments on spontaneous English speech show that our technique yields a relative WER improvement up to 9.87% for utterances by female speakers, with larger gains for the least-represented f0 ranges.
翻译:自动语音识别(ASR)系统对语音数据的社会语言学变异性较为敏感,其中性别因素扮演着关键角色。由于训练数据中女性语音的代表性不足,这可能导致男性和女性说话者在识别准确率上存在差异。尽管在混合ASR模型背景下已提出多种解决方案,但端到端神经架构中的性别偏差问题尚未得到明确解决。为填补这一空白,我们提出一种数据增强技术,通过操控基频(f0)和共振峰来模拟代表性不足的女性语音,从而减少性别间的数据不平衡并增加各性别组内的变异性。在自发性英语语音上的实验表明,该技术对女性说话者的语音实现了高达9.87%的相对词错误率改善,在代表性最低的f0范围内获得了更大增益。