Recent advances in eXplainable AI (XAI) have provided new insights into how models for vision, language, and tabular data operate. However, few approaches exist for understanding speech models. Existing work focuses on a few spoken language understanding (SLU) tasks, and explanations are difficult to interpret for most users. We introduce a new approach to explain speech classification models. We generate easy-to-interpret explanations via input perturbation on two information levels. 1) Word-level explanations reveal how each word-related audio segment impacts the outcome. 2) Paralinguistic features (e.g., prosody and background noise) answer the counterfactual: ``What would the model prediction be if we edited the audio signal in this way?'' We validate our approach by explaining two state-of-the-art SLU models on two speech classification tasks in English and Italian. Our findings demonstrate that the explanations are faithful to the model's inner workings and plausible to humans. Our method and findings pave the way for future research on interpreting speech models.
翻译:可解释人工智能(XAI)的最新进展为视觉、语言和表格数据模型的工作原理提供了新见解。然而,针对语音模型的理解方法仍然较少。现有工作主要集中于少数口语理解(SLU)任务,且其解释对大多数用户而言难以理解。我们提出了一种新的语音分类模型解释方法。通过两个信息层面的输入扰动,生成易于理解的解释:1) 词级解释揭示每个与词语相关的音频片段如何影响输出结果;2) 副语言特征(例如韵律和背景噪声)回答反事实问题:“若以这种方式编辑音频信号,模型预测结果会怎样?”我们通过在英语和意大利语的两种语音分类任务上解释两个最先进的SLU模型来验证所提方法。结果表明,这些解释对模型内部机制具有忠实性,且对人类而言具有合理性。我们的方法与发现为未来语音模型解释研究奠定了基础。