Self-supervised learning (SSL) leverages large datasets of unlabeled speech to reach impressive performance with reduced amounts of annotated data. The high number of proposed approaches fostered the emergence of comprehensive benchmarks that evaluate their performance on a set of downstream tasks exploring various aspects of the speech signal. However, while the number of considered tasks has been growing, most proposals rely upon a single downstream architecture that maps the frozen SSL representations to the task labels. This study examines how benchmarking results are affected by changes in the probing head architecture. Interestingly, we found that altering the downstream architecture structure leads to significant fluctuations in the performance ranking of the evaluated models. Against common practices in speech SSL benchmarking, we evaluate larger-capacity probing heads, showing their impact on performance, inference costs, generalization and multi-level feature exploitation.
翻译:自监督学习利用大量无标注语音数据,在减少标注数据量的同时取得了显著性能。由于提出的方法众多,催生了综合性基准测试的出现,这些测试通过评估模型在一系列探索语音信号不同方面的下游任务上的表现来比较其性能。然而,尽管被考虑的任务数量不断增加,大多数方案仍依赖于单一的下游架构,该架构将冻结的自监督表示映射到任务标签上。本研究探讨了探测头架构的变化对基准测试结果的影响。有趣的是,我们发现改变下游架构结构会导致被评估模型性能排名的显著波动。与语音自监督基准测试中的常见做法不同,我们评估了更大容量的探测头,展示了它们在性能、推理成本、泛化能力和多层级特征利用方面的影响。