The challenge of fairness arises when Automatic Speech Recognition (ASR) systems do not perform equally well for all sub-groups of the population. In the past few years there have been many improvements in overall speech recognition quality, but without any particular focus on advancing Equality and Equity for all user groups for whom systems do not perform well. ASR fairness is therefore also a robustness issue. Meanwhile, data privacy also takes priority in production systems. In this paper, we present a privacy preserving approach to improve fairness and robustness of end-to-end ASR without using metadata, zip codes, or even speaker or utterance embeddings directly in training. We extract utterance level embeddings using a speaker ID model trained on a public dataset, which we then use in an unsupervised fashion to create acoustic clusters. We use cluster IDs instead of speaker utterance embeddings as extra features during model training, which shows improvements for all demographic groups and in particular for different accents.
翻译:自动语音识别系统在不同人群子组上表现不一致时,就会产生公平性挑战。近年来,尽管整体语音识别质量取得了显著进步,但尚未特别关注提升系统对表现欠佳用户群体的平等与公正性。因此,ASR公平性本质上也是鲁棒性问题。同时,数据隐私在生产系统中同样占据优先地位。本文提出一种隐私保护方法,在不使用元数据、邮政编码、说话人嵌入或话语嵌入直接参与训练的情况下,提升端到端ASR的公平性与鲁棒性。我们利用在公开数据集上训练的说话人身份模型提取话语级嵌入,随后通过无监督方式创建声学聚类。在模型训练过程中,我们使用聚类标识代替说话人话语嵌入作为附加特征,实验表明该方法能改善所有人口群体的表现,尤其对不同口音群体效果显著。