We present ProsAudit, a benchmark in English to assess structural prosodic knowledge in self-supervised learning (SSL) speech models. It consists of two subtasks, their corresponding metrics, and an evaluation dataset. In the protosyntax task, the model must correctly identify strong versus weak prosodic boundaries. In the lexical task, the model needs to correctly distinguish between pauses inserted between words and within words. We also provide human evaluation scores on this benchmark. We evaluated a series of SSL models and found that they were all able to perform above chance on both tasks, even when evaluated on an unseen language. However, non-native models performed significantly worse than native ones on the lexical task, highlighting the importance of lexical knowledge in this task. We also found a clear effect of size with models trained on more data performing better in the two subtasks.
翻译:本文提出ProsAudit——一个用于评估自监督学习(SSL)语音模型中结构性韵律知识的英文基准测试。该基准包含两个子任务、对应评价指标及评估数据集。在原型句法任务中,模型需正确识别强/弱韵律边界;在词汇任务中,模型需准确区分词间停顿与词内停顿。我们同时提供了该基准上的人类评估分数。通过对一系列SSL模型进行评估发现,所有模型均能在两项任务中取得优于随机水平的性能,即使对未见语言进行评估时亦是如此。然而,在词汇任务中,非母语模型的性能显著低于母语模型,凸显了该任务中词汇知识的重要性。我们还观察到数据规模对性能的显著影响——使用更丰富数据训练的模型在两个子任务中表现更优。