In this study, formant tracking is investigated by refining the formants tracked by an existing data-driven tracker, DeepFormants, using the formants estimated in a model-driven manner by linear prediction (LP)-based methods. As LP-based formant estimation methods, conventional covariance analysis (LP-COV) and the recently proposed quasi-closed phase forward-backward (QCP-FB) analysis are used. In the proposed refinement approach, the contours of the three lowest formants are first predicted by the data-driven DeepFormants tracker, and the predicted formants are replaced frame-wise with local spectral peaks shown by the model-driven LP-based methods. The refinement procedure can be plugged into the DeepFormants tracker with no need for any new data learning. Two refined DeepFormants trackers were compared with the original DeepFormants and with five known traditional trackers using the popular vocal tract resonance (VTR) corpus. The results indicated that the data-driven DeepFormants trackers outperformed the conventional trackers and that the best performance was obtained by refining the formants predicted by DeepFormants using QCP-FB analysis. In addition, by tracking formants using VTR speech that was corrupted by additive noise, the study showed that the refined DeepFormants trackers were more resilient to noise than the reference trackers. In general, these results suggest that LP-based model-driven approaches, which have traditionally been used in formant estimation, can be combined with a modern data-driven tracker easily with no further training to improve the tracker's performance.
翻译:本研究通过结合基于线性预测(LP)的模型驱动方法所估计的共振峰,对现有数据驱动跟踪器DeepFormants的共振峰追踪结果进行精炼。采用的LP共振峰估计方法包括传统协方差分析(LP-COV)与近期提出的准闭相前向-后向分析(QCP-FB)。在所提出的精炼方法中,首先由数据驱动的DeepFormants跟踪器预测三个最低共振峰的轨迹,随后逐帧将预测的共振峰替换为模型驱动方法所显示的局部频谱峰值。该精炼过程可直接嵌入DeepFormants跟踪器,无需额外数据学习。将两种精炼后的DeepFormants跟踪器与原始DeepFormants及五种已知传统跟踪器进行对比,采用广泛使用的声道共振(VTR)语料库。结果表明,数据驱动的DeepFormants跟踪器优于传统跟踪器,其中通过QCP-FB分析精炼DeepFormants预测共振峰的方法获得最佳性能。此外,在受加性噪声污染的VTR语音共振峰追踪实验中,精炼后的DeepFormants跟踪器显示出比参考跟踪器更强的抗噪性。总体而言,这些结果表明,传统用于共振峰估计的基于LP的模型驱动方法可轻松与现代数据驱动跟踪器结合,无需额外训练即可提升跟踪性能。