Many self-supervised speech models (S3Ms) have been introduced over the last few years, improving performance and data efficiency on various speech tasks. However, these empirical successes alone do not give a complete picture of what is learned during pre-training. Recent work has begun analyzing how S3Ms encode certain properties, such as phonetic and speaker information, but we still lack a proper understanding of knowledge encoded at the word level and beyond. In this work, we use lightweight analysis methods to study segment-level linguistic properties -- word identity, boundaries, pronunciation, syntactic features, and semantic features -- encoded in S3Ms. We present a comparative study of layer-wise representations from ten S3Ms and find that (i) the frame-level representations within each word segment are not all equally informative, and (ii) the pre-training objective and model size heavily influence the accessibility and distribution of linguistic information across layers. We also find that on several tasks -- word discrimination, word segmentation, and semantic sentence similarity -- S3Ms trained with visual grounding outperform their speech-only counterparts. Finally, our task-based analyses demonstrate improved performance on word segmentation and acoustic word discrimination while using simpler methods than prior work.
翻译:近年来,多种自监督语音模型(S3Ms)被提出,在各类语音任务中提升了性能和数据效率。然而,这些经验性成功本身并未完整揭示预训练过程中模型所学的内容。已有研究开始分析S3Ms如何编码语音性和说话人信息等特定属性,但我们仍缺乏对词汇层面及以上知识编码情况的充分理解。本研究采用轻量分析方法,探究S3Ms编码的片段级语言属性——词汇身份、边界、发音、句法特征和语义特征。我们对十个S3Ms的逐层表示进行对比研究发现:(i)每个词汇片段内的帧级表示并非均等包含信息;(ii)预训练目标和模型规模会显著影响语言信息在各层中的可访问性与分布。此外,在词汇区分、词汇分割和语义句子相似性等多项任务中,基于视觉锚定的S3Ms性能优于纯语音模型。最后,我们的任务分析表明,相较于先前研究使用更简单的方法,在词汇分割和声学词汇区分任务上取得了更优效果。