Depression is a common mental disorder. Automatic depression detection tools using speech, enabled by machine learning, help early screening of depression. This paper addresses two limitations that may hinder the clinical implementations of such tools: noise resulting from segment-level labelling and a lack of model interpretability. We propose a bi-modal speech-level transformer to avoid segment-level labelling and introduce a hierarchical interpretation approach to provide both speech-level and sentence-level interpretations, based on gradient-weighted attention maps derived from all attention layers to track interactions between input features. We show that the proposed model outperforms a model that learns at a segment level ($p$=0.854, $r$=0.947, $F1$=0.947 compared to $p$=0.732, $r$=0.808, $F1$=0.768). For model interpretation, using one true positive sample, we show which sentences within a given speech are most relevant to depression detection; and which text tokens and Mel-spectrogram regions within these sentences are most relevant to depression detection. These interpretations allow clinicians to verify the validity of predictions made by depression detection tools, promoting their clinical implementations.
翻译:抑郁症是一种常见的精神障碍。基于机器学习、利用语音的自动抑郁症检测工具有助于抑郁症的早期筛查。本文针对可能阻碍此类工具临床实施的两个局限性:分段级别标注产生的噪声以及模型可解释性的缺乏。我们提出了一种双模态语音级Transformer以规避分段级别标注,并引入了一种层级解释方法,基于从所有注意力层提取的梯度加权注意力图来追踪输入特征之间的交互,从而提供语音级和句子级两种层级的解释。我们证明,所提出的模型优于在分段级别学习的模型(精度$p$=0.854,召回率$r$=0.947,$F1$分数=0.947,相比之下对比模型为$p$=0.732,$r$=0.808,$F1$=0.768)。在模型解释方面,通过一个真阳性样本,我们展示了给定语音中哪些句子与抑郁症检测最相关;以及这些句子中哪些文本标记和梅尔频谱图区域与抑郁症检测最相关。这些解释使临床医生能够验证抑郁症检测工具所作预测的有效性,从而推动其临床实施。