Speech Emotion Recognition (SER) is a critical enabler of emotion-aware communication in human-computer interactions. Recent advancements in Deep Learning (DL) have substantially enhanced the performance of SER models through increased model complexity. However, designing optimal DL architectures requires prior experience and experimental evaluations. Encouragingly, Neural Architecture Search (NAS) offers a promising avenue to determine an optimal DL model automatically. In particular, Differentiable Architecture Search (DARTS) is an efficient method of using NAS to search for optimised models. This paper proposes a DARTS-optimised joint CNN and LSTM architecture, to improve SER performance, where the literature informs the selection of CNN and LSTM coupling to offer improved performance. While DARTS has previously been applied to CNN and LSTM combinations, our approach introduces a novel mechanism, particularly in selecting CNN operations using DARTS. In contrast to previous studies, we refrain from imposing constraints on the order of the layers for the CNN within the DARTS cell; instead, we allow DARTS to determine the optimal layer order autonomously. Experimenting with the IEMOCAP and MSP-IMPROV datasets, we demonstrate that our proposed methodology achieves significantly higher SER accuracy than hand-engineering the CNN-LSTM configuration. It also outperforms the best-reported SER results achieved using DARTS on CNN-LSTM.
翻译:语音情感识别是人机交互中实现情感感知通信的关键技术。近年来,深度学习通过增加模型复杂度显著提升了语音情感识别模型的性能。然而,设计最优的深度学习架构需要先验知识和实验评估。值得关注的是,神经架构搜索为自动确定最优深度学习模型提供了有效途径。其中,可微分架构搜索是利用神经架构搜索寻找优化模型的高效方法。本文提出一种可微分架构搜索优化的卷积神经网络与长短期记忆网络联合架构,用于提升语音情感识别性能。现有文献表明卷积神经网络与长短期记忆网络的耦合能提供更优性能。虽然可微分架构搜索此前已被应用于卷积神经网络与长短期记忆网络的组合,但我们的方法引入了一种创新机制,特别是在使用可微分架构搜索选择卷积神经网络运算方面。与以往研究不同,我们不对可微分架构搜索单元内卷积神经网络的层序施加约束,而是让可微分架构搜索自主确定最优层序。通过在IEMOCAP和MSP-IMPROV数据集上的实验证明,我们提出的方法相比人工设计的卷积神经网络-长短期记忆网络配置,能实现显著更高的语音情感识别准确率,并且在卷积神经网络-长短期记忆网络上达到了优于现有可微分架构搜索语音情感识别最佳记录的性能。