Within the area of speech enhancement, there is an ongoing interest in the creation of neural systems which explicitly aim to improve the perceptual quality of the processed audio. In concert with this is the topic of non-intrusive (i.e. without clean reference) speech quality prediction, for which neural networks are trained to predict human-assigned quality labels directly from distorted audio. When combined, these areas allow for the creation of powerful new speech enhancement systems which can leverage large real-world datasets of distorted audio, by taking inference of a pre-trained speech quality predictor as the sole loss function of the speech enhancement system. This paper aims to identify a potential pitfall with this approach, namely hallucinations which are introduced by the enhancement system `tricking' the speech quality predictor.
翻译:在语音增强领域,始终存在对明确旨在提升处理音频感知质量的神经系统的研究兴趣。与此相关的是非侵入式(即无需纯净参考信号)语音质量预测课题——神经网络通过训练直接从失真音频中预测人类标注的质量标签。将这两者结合,可通过将预训练语音质量预测器的推理结果作为语音增强系统的唯一损失函数,构建能利用大规模真实失真音频数据集的新型强效语音增强系统。本文旨在揭示该方法的一个潜在缺陷,即增强系统通过"欺骗"语音质量预测器引入的幻觉现象。