Useful conversational agents must accurately capture named entities to minimize error for downstream tasks, for example, asking a voice assistant to play a track from a certain artist, initiating navigation to a specific location, or documenting a laboratory result for a patient. However, where named entities such as ``Ukachukwu`` (Igbo), ``Lakicia`` (Swahili), or ``Ingabire`` (Rwandan) are spoken, automatic speech recognition (ASR) models' performance degrades significantly, propagating errors to downstream systems. We model this problem as a distribution shift and demonstrate that such model bias can be mitigated through multilingual pre-training, intelligent data augmentation strategies to increase the representation of African-named entities, and fine-tuning multilingual ASR models on multiple African accents. The resulting fine-tuned models show an 81.5\% relative WER improvement compared with the baseline on samples with African-named entities.
翻译:实用的对话代理必须准确捕获命名实体,以最小化下游任务的错误,例如要求语音助手播放某位艺术家的曲目、导航至特定地点或记录患者的化验结果。然而,当“Ukachukwu”(伊博语)、“Lakicia”(斯瓦希里语)或“Ingabire”(卢旺达语)等命名实体被说出时,自动语音识别(ASR)模型的性能显著下降,并将错误传播至下游系统。我们将此问题建模为分布偏移,并证明这种模型偏差可通过多语言预训练、增加非洲命名实体表示比例的智能数据增强策略,以及在多种非洲口音上微调多语言ASR模型来缓解。最终微调后的模型在包含非洲命名实体的样本上,与基线相比实现了81.5%的相对词错误率(WER)改善。