Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. However, the quality of these models is often measured by the performance of different downstream tasks. How well the representations access the information of interest is less studied. In this work, we take a closer look into existing self-supervised methods of speech from an information-theoretic perspective. We aim to develop metrics using mutual information to help practical problems such as model design and selection. We use linear probes to estimate the mutual information between the target information and learned representations, showing another insight into the accessibility to the target information from speech representations. Further, we explore the potential of evaluating representations in a self-supervised fashion, where we estimate the mutual information between different parts of the data without using any labels. Finally, we show that both supervised and unsupervised measures echo the performance of the models on layer-wise linear probing and speech recognition.
翻译:现有关于自监督语音表征学习的研究主要集中于开发新的训练方法,并将预训练模型应用于不同场景。然而,这些模型的质量通常通过不同下游任务的表现来衡量,而表征对目标信息的可访问性研究则相对不足。本研究从信息论视角深入剖析现有自监督语音方法,旨在利用互信息开发评估指标以辅助模型设计与选择等实际问题。通过线性探针估算目标信息与学习表征之间的互信息,揭示了语音表征中目标信息可访问性的新视角。进一步地,我们探索了以自监督方式评估表征的潜力——无需使用任何标签即可估算数据不同部分之间的互信息。最后,研究表明,有监督与无监督评估指标均与模型在逐层线性探针及语音识别任务上的表现相吻合。