Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency representations, e.g, the log-Mel spectrogram's average or maximum over time, do not always work well for different machines. This paper presents Time-Weighted Frequency Domain Representation (TWFR) with the GMM method (TWFR-GMM) for anomalous sound detection. The TWFR is a generalized statistical frequency domain representation that can adapt to different machine types, using the global weighted ranking pooling over time-domain. This allows GMM estimator to recognize anomalies, even under domain-shift conditions, as visualized with a Mahalanobis distance-based metric. Experiments on DCASE 2022 Challenge Task2 dataset show that our method has better detection performance than recent deep learning methods. TWFR-GMM is the core of our submission that achieved the 3rd place in DCASE 2022 Challenge Task2.
翻译:尽管深度学习是无监督异常声音检测的主流方法,但以统计音频频率表示作为输入的混合高斯模型(GMM)能够以更低的模型复杂度和更少的参数达到可比较的结果。现有的统计频率表示方法(例如,对数梅尔频谱图的时间平均或最大值)对不同机器并不总是有效。本文提出了基于GMM方法的时域加权频率表示(TWFR-GMM)用于异常声音检测。TWFR是一种广义的统计频率域表示方法,通过时域全局加权排序池化技术,能够适应不同的机器类型。这使得GMM估计器即使在域迁移条件下也能识别异常,并通过基于马氏距离的度量进行可视化。在DCASE 2022挑战赛Task2数据集上的实验表明,我们的方法比近期深度学习方法具有更好的检测性能。TWFR-GMM是我们获得DCASE 2022挑战赛Task2第三名解决方案的核心。