This work profoundly analyzes discrete self-supervised speech representations through the eyes of Generative Spoken Language Modeling (GSLM). Following the findings of such an analysis, we propose practical improvements to the discrete unit for the GSLM. First, we start comprehending these units by analyzing them in three axes: interpretation, visualization, and resynthesis. Our analysis finds a high correlation between the speech units to phonemes and phoneme families, while their correlation with speaker or gender is weaker. Additionally, we found redundancies in the extracted units and claim that one reason may be the units' context. Following this analysis, we propose a new, unsupervised metric to measure unit redundancies. Finally, we use this metric to develop new methods that improve the robustness of units clustering and show significant improvement considering zero-resource speech metrics such as ABX. Code and analysis tools are available under the following link.
翻译:本工作从生成式口语语言建模(GSLM)的视角深入分析了离散自监督语音表示。基于此类分析的结果,我们提出了针对GSLM离散单元的实用改进方案。首先,我们沿着三个维度——解释、可视化和再合成——对这些单元进行理解性分析。分析发现,语音单元与音素及音素家族之间存在高度相关性,而与说话人或性别之间的相关性较弱。此外,我们观察到提取的单元中存在冗余,并推测其原因之一可能来自单元的上下文。基于这一分析,我们提出了一种新的无监督度量指标以衡量单元冗余度。最终,我们利用该指标开发了新方法,改进了单元聚类的稳健性,并在ABX等零资源语音评估指标上取得了显著提升。相关代码与分析工具可通过以下链接获取。