Due to the absence of explicit word boundaries in the speech stream, the task of segmenting spoken sentences into word units without text supervision is particularly challenging. In this work, we leverage the most recent self-supervised speech models that have proved to quickly adapt to new tasks through fine-tuning, even in low resource conditions. Taking inspiration from semi-supervised learning, we fine-tune an XLS-R model to predict word boundaries themselves produced by top-tier speech segmentation systems: DPDP, VG-HuBERT, GradSeg and DP-Parse. Once XLS-R is fine-tuned, it is used to infer new word boundary labels that are used in turn for another fine-tuning step. Our method consistently improves the performance of each system and sets a new state-of-the-art that is, on average 130% higher than the previous one as measured by the F1 score on correctly discovered word tokens on five corpora featuring different languages. Finally, our system can segment speech from languages unseen during fine-tuning in a zero-shot fashion.
翻译:由于语音流中缺乏显式词边界,在没有文本监督的情况下将语音句子分割为词单元的任务尤为困难。本研究利用最新的自监督语音模型,这些模型通过微调能够快速适应新任务,即使在低资源条件下也是如此。受半监督学习启发,我们对XLS-R模型进行微调,使其能够预测由顶级语音分割系统(DPDP、VG-HuBERT、GradSeg和DP-Parse)生成的词边界。一旦XLS-R完成微调,它被用于推断新的词边界标签,这些标签随后被用于另一轮微调。我们的方法持续提升了每个系统的性能,并设立了新的最高水平,在五个涵盖不同语言的语料库上,以正确发现的词令牌的F1分数衡量,平均比先前水平高出130%。最后,我们的系统能够以零样本方式对微调中未见语言的语音进行分割。