This paper proposes handling training data sparsity in speech-based automatic depression detection (SDD) using foundation models pre-trained with self-supervised learning (SSL). An analysis of SSL representations derived from different layers of pre-trained foundation models is first presented for SDD, which provides insight to suitable indicator for depression detection. Knowledge transfer is then performed from automatic speech recognition (ASR) and emotion recognition to SDD by fine-tuning the foundation models. Results show that the uses of oracle and ASR transcriptions yield similar SDD performance when the hidden representations of the ASR model is incorporated along with the ASR textual information. By integrating representations from multiple foundation models, state-of-the-art SDD results based on real ASR were achieved on the DAIC-WOZ dataset.
翻译:本文提出利用基于自监督学习(SSL)预训练的基座模型处理语音自动抑郁检测(SDD)中训练数据稀疏的问题。首先分析了预训练基座模型不同层中提取的SSL表示对SDD的作用,为抑郁检测的合适指标提供了见解。随后通过微调基座模型,将自动语音识别(ASR)和情感识别的知识迁移至SDD。结果表明,当ASR模型的隐表示与ASR文本信息结合使用时,使用真实转录和ASR转录的SDD性能相近。通过整合多个基座模型的表示,在DAIC-WOZ数据集上基于真实ASR实现了最先进的SDD性能。