The sequence length along the time axis is often the dominant factor of the computation in speech processing. Works have been proposed to reduce the sequence length for lowering the computational cost in self-supervised speech models. However, different downstream tasks have different tolerance of sequence compressing, so a model that produces a fixed compressing rate may not fit all tasks. In this work, we introduce a once-for-all (OFA) sequence compression framework for self-supervised speech models that supports a continuous range of operating compressing rates. The framework is evaluated on various tasks, showing marginal degradation compared to the fixed compressing rate variants with a smooth performance-efficiency trade-off. We further explore adaptive compressing rate learning, demonstrating the ability to select task-specific preferred frame periods without needing a grid search.
翻译:时间轴上的序列长度通常是语音处理中计算量的主导因素。已有研究提出通过缩短序列长度来降低自监督语音模型的计算成本。然而,不同下游任务对序列压缩的容忍度存在差异,因此固定压缩率的模型可能无法适应所有任务。本文提出了一种用于自监督语音模型的一次性序列压缩框架,该框架支持连续范围内的可变操作压缩率。通过在多种任务上的评估,该框架相较于固定压缩率变体仅产生轻微性能退化,同时实现平滑的性能-效率权衡。我们进一步探索了自适应压缩率学习方法,展示了无需网格搜索即可选择任务特定优选帧周期的能力。