It has been generally assumed in the automatic speech recognition (ASR) literature that it is better for models to have access to wider context windows. Yet, many of the potential reasons this might be true in the supervised setting do not necessarily transfer over to the case of unsupervised learning. We investigate how much context is necessary to achieve high-quality pre-trained acoustic models using self-supervised learning. We principally investigate contrastive predictive coding (CPC), which we adapt to be able to precisely control the amount of context visible to the model during training and inference. We find that phone discriminability in the resulting model representations peaks at around 40~ms of preceding context, and that having too much context (beyond around 320 ms) substantially degrades the quality of the representations. Surprisingly, we find that this pattern also transfers to supervised ASR when the pre-trained representations are used as frozen input features. Our results point to potential changes in the design of current upstream architectures to better facilitate a variety of downstream tasks.
翻译:在自动语音识别(ASR)文献中,普遍认为模型拥有更宽的上下文窗口会更好。然而,许多在监督学习场景下支持这一假设的潜在原因,并不一定适用于无监督学习的情况。我们研究了通过自监督学习获得高质量预训练声学模型所需的上文信息量。主要研究了对比预测编码(CPC),并对其进行了改进,以便能在训练和推理过程中精确控制模型可见的上文信息量。我们发现,所得到的模型表征中的音素可辨别性在大约40毫秒的前置上下文处达到峰值,而过多的上下文(超过约320毫秒)会显著降低表征质量。令人惊讶的是,当预训练表征作为冻结的输入特征用于监督式ASR时,这一模式同样适用。我们的研究结果提示,当前上游架构的设计可能需要改变,以更好地支持各种下游任务。