Hidden-unit BERT (HuBERT) is a widely-used self-supervised learning (SSL) model in speech processing. However, we argue that its fixed 20ms resolution for hidden representations would not be optimal for various speech-processing tasks since their attributes (e.g., speaker characteristics and semantics) are based on different time scales. To address this limitation, we propose utilizing HuBERT representations at multiple resolutions for downstream tasks. We explore two approaches, namely the parallel and hierarchical approaches, for integrating HuBERT features with different resolutions. Through experiments, we demonstrate that HuBERT with multiple resolutions outperforms the original model. This highlights the potential of utilizing multiple resolutions in SSL models like HuBERT to capture diverse information from speech signals.
翻译:隐单元BERT(HuBERT)是一种广泛应用于语音处理的自监督学习模型。然而,我们认为其固定的20ms隐表示分辨率未必适用于各类语音处理任务,因为不同任务属性(如说话人特征与语义信息)所依据的时间尺度不同。为解决这一局限性,我们提出在下游任务中采用多分辨率的HuBERT表示。我们探索了两种方法——并行法与层级法——用于整合不同分辨率的HuBERT特征。通过实验证明,采用多分辨率的HuBERT模型性能优于原始模型。这凸显了在HuBERT等自监督学习模型中利用多分辨率机制,以从语音信号中捕获多样化信息的潜力。