Deep neural networks are a promising tool for Audio Event Classification. In contrast to other data like natural images, there are many sensible and non-obvious representations for audio data, which could serve as input to these models. Due to their black-box nature, the effect of different input representations has so far mostly been investigated by measuring classification performance. In this work, we leverage eXplainable AI (XAI), to understand the underlying classification strategies of models trained on different input representations. Specifically, we compare two model architectures with regard to relevant input features used for Audio Event Detection: one directly processes the signal as the raw waveform, and the other takes in its time-frequency spectrogram representation. We show how relevance heatmaps obtained via "Siren"{Layer-wise Relevance Propagation} uncover representation-dependent decision strategies. With these insights, we can make a well-informed decision about the best input representation in terms of robustness and representativity and confirm that the model's classification strategies align with human requirements.
翻译:深度神经网络是音频事件分类领域极具前景的工具。与自然图像等其他数据类型不同,音频数据存在多种合理且非显而易见的表示形式可作为模型输入。由于深度神经网络的"黑箱"特性,不同输入表示的影响此前主要通过分类性能指标进行研究。本研究利用可解释人工智能(XAI)技术,理解基于不同输入表示所训练模型的潜在分类策略。具体而言,我们对比了两种模型架构在音频事件检测中使用的相关输入特征:一种直接以原始波形处理信号,另一种则采用其时频频谱图表示。通过"Siren"{逐层相关性传播}方法生成的相关性热力图,我们揭示了依赖于表示形式的决策策略。基于这些见解,我们能够在鲁棒性和代表性方面做出关于最佳输入表示的明智决策,并证实模型的分类策略与人类需求保持一致。