Personal Digital Assistants (PDAs) - such as Siri, Alexa and Google Assistant, to name a few - play an increasingly important role to access information and complete tasks spanning multiple domains, and by diverse groups of users. A text-to-speech (TTS) module allows PDAs to interact in a natural, human-like manner, and play a vital role when the interaction involves people with visual impairments or other disabilities. To cater to the needs of a diverse set of users, inclusive TTS is important to recognize and pronounce correctly text in different languages and dialects. Despite great progress in speech synthesis, the pronunciation accuracy of named entities in a multi-lingual setting still has a large room for improvement. Existing approaches to correct named entity (NE) mispronunciations, like retraining Grapheme-to-Phoneme (G2P) models, or maintaining a TTS pronunciation dictionary, require expensive annotation of the ground truth pronunciation, which is also time consuming. In this work, we present a highly-precise, PDA-compatible pronunciation learning framework for the task of TTS mispronunciation detection and correction. In addition, we also propose a novel mispronunciation detection model called DTW-SiameseNet, which employs metric learning with a Siamese architecture for Dynamic Time Warping (DTW) with triplet loss. We demonstrate that a locale-agnostic, privacy-preserving solution to the problem of TTS mispronunciation detection is feasible. We evaluate our approach on a real-world dataset, and a corpus of NE pronunciations of an anonymized audio dataset of person names recorded by participants from 10 different locales. Human evaluation shows our proposed approach improves pronunciation accuracy on average by ~6% compared to strong phoneme-based and audio-based baselines.
翻译:个人数字助理(PDA),如Siri、Alexa和Google Assistant等,在跨多个领域的信息获取与任务完成中发挥着日益重要的作用,并服务于多元化的用户群体。文本转语音(TTS)模块使PDA能够以自然、类人的方式进行交互,尤其在涉及视障或其他残障人士的交互中扮演关键角色。为满足多元化用户需求,包容性TTS需能正确识别并准确发音不同语言及方言中的文本。尽管语音合成已取得巨大进展,但在多语言场景下命名实体的发音准确性仍有较大提升空间。现有纠正命名实体(NE)发音错误的方法,如重新训练字素到音素(G2P)模型或维护TTS发音词典,均需昂贵且耗时的真实发音标注。本研究针对TTS误发音检测与纠正任务,提出了一种高精度、兼容PDA的发音学习框架。同时,我们提出了一种名为DTW-SiameseNet的新型误发音检测模型,该模型采用孪生架构的度量学习,结合三元组损失实现动态时间规整(DTW)。我们证明,面向TTS误发音检测问题的与区域无关且隐私保护的解决方案是可行的。我们在真实数据集及一个匿名音频数据集(包含来自10个不同地区的参与者录制的人名NE发音语料库)上评估了该方法。人工评估表明,与基于音素及音频的强基线方法相比,所提方法平均将发音准确性提升了约6%。