We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, finetuned on speaker identification (SID), finetuned on automatic speech recognition (ASR), and ASR-finetuned using a fairness enhancing algorithm. We find that S3Ms encode information about several speaker group categories (SGCs), including their gender, age, dialect, ethnicity, and whether they are a native speaker. We find that finetuning for SID amplifies certain SGCs, namely those whose variance is more phonetic in nature, though it does not amplify other SGCs, namely those whose variance is more semantic in nature. On the other hand, finetuning for ASR discards phonetically variant speaker group information (SGI) but retains semantically variant SGI. We find that ASR algorithms designed for fairness improvement change to what extent SGI is encoded in S3Ms; however, this is primarily true for for phonetically variant SGCs, and less true for semantically variant SGCs. We discuss how SGI is encoded by each layer, and identify subdimensions of embeddings responsible for encoding different SGCs. Finally, we discuss how our findings could be beneficial in designing fairer ASR algorithms.
翻译:我们研究了自监督语音识别模型(S3Ms)对说话人群体(SGs)的学习情况。我们考察了S3Ms的几种状态:预训练、针对说话人识别(SID)微调、针对自动语音识别(ASR)微调,以及使用公平性增强算法进行ASR微调。我们发现,S3Ms编码了关于多个说话人群体类别(SGCs)的信息,包括其性别、年龄、方言、种族以及是否为母语者。我们发现,针对SID的微调放大了某些SGCs,即那些方差更偏向语音学性质的SGCs,但并未放大其他方差更偏向语义学性质的SGCs。另一方面,针对ASR的微调丢弃了语音学变异的说话人群体信息(SGI),但保留了语义学变异的SGI。我们发现,旨在提升公平性的ASR算法会改变S3Ms中编码的SGI程度;然而,这主要适用于语音学变异的SGCs,而对语义学变异的SGCs影响较小。我们讨论了各层如何编码SGI,并识别了嵌入中负责编码不同SGCs的子维度。最后,我们讨论了研究结果如何有助于设计更公平的ASR算法。