Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multitude of accents, particularly in the African context, remains impractical due to their sheer diversity and associated budget constraints. To address these challenges, we propose \textit{AccentFold}, a method that exploits spatial relationships between learned accent embeddings to improve downstream Automatic Speech Recognition (ASR). Our exploratory analysis of speech embeddings representing 100+ African accents reveals interesting spatial accent relationships highlighting geographic and genealogical similarities, capturing consistent phonological, and morphological regularities, all learned empirically from speech. Furthermore, we discover accent relationships previously uncharacterized by the Ethnologue. Through empirical evaluation, we demonstrate the effectiveness of AccentFold by showing that, for out-of-distribution (OOD) accents, sampling accent subsets for training based on AccentFold information outperforms strong baselines a relative WER improvement of 4.6%. AccentFold presents a promising approach for improving ASR performance on accented speech, particularly in the context of African accents, where data scarcity and budget constraints pose significant challenges. Our findings emphasize the potential of leveraging linguistic relationships to improve zero-shot ASR adaptation to target accents.
翻译:尽管语音识别技术取得了进步,带口音语音仍具挑战性。以往方法侧重于建模技术或构建带口音语音数据集,但由于非洲口音种类繁多且预算有限,为众多口音收集充足数据并不现实。为应对这些挑战,我们提出AccentFold方法,利用学习到的口音嵌入之间的空间关系来改进下游自动语音识别(ASR)。我们对代表100多种非洲口音的语音嵌入进行探索性分析,揭示了有趣的空间口音关系,这些关系体现了地理和谱系相似性,捕捉了语音中经验习得的、一致的音系和形态学规律。此外,我们还发现了《民族语》此前未表征的口音关系。通过实证评估,我们证明了AccentFold的有效性:对于分布外(OOD)口音,基于AccentFold信息采样口音子集进行训练,相比强基线实现了4.6%的相对词错误率(WER)改进。AccentFold为提升带口音语音的ASR性能提供了有前景的方法,尤其在数据稀缺和预算受限的非洲口音场景中。我们的研究结果凸显了利用语言关系改善零样本ASR目标口音适配的潜力。