Streaming models are an essential component of real-time speech enhancement tools. The streaming regime constrains speech enhancement models to use only a tiny context of future information. As a result, the low-latency streaming setup is generally considered a challenging task and has a significant negative impact on the model's quality. However, the sequential nature of streaming generation offers a natural possibility for autoregression, that is, utilizing previous predictions while making current ones. The conventional method for training autoregressive models is teacher forcing, but its primary drawback lies in the training-inference mismatch that can lead to a substantial degradation in quality. In this study, we propose a straightforward yet effective alternative technique for training autoregressive low-latency speech enhancement models. We demonstrate that the proposed approach leads to stable improvement across diverse architectures and training scenarios.
翻译:流式模型是实时语音增强工具的关键组成部分。流式工作模式限制了语音增强模型仅能使用极少的未来信息上下文,因此低延迟流式任务通常被视为具有挑战性的问题,并对模型质量产生显著的负面影响。然而,流式生成的序列特性为自回归提供了天然可能性,即在当前预测中利用先前的预测结果。训练自回归模型的传统方法是教师强制,但其主要缺陷在于训练与推理之间的不匹配,这可能导致质量严重下降。本研究提出了一种简单而有效的替代技术,用于训练自回归低延迟语音增强模型。我们证明,所提出的方法能在不同架构和训练场景下实现稳定的性能提升。