Effective speech emotional representations play a key role in Speech Emotion Recognition (SER) and Emotional Text-To-Speech (TTS) tasks. However, emotional speech samples are more difficult and expensive to acquire compared with Neutral style speech, which causes one issue that most related works unfortunately neglect: imbalanced datasets. Models might overfit to the majority Neutral class and fail to produce robust and effective emotional representations. In this paper, we propose an Emotion Extractor to address this issue. We use augmentation approaches to train the model and enable it to extract effective and generalizable emotional representations from imbalanced datasets. Our empirical results show that (1) for the SER task, the proposed Emotion Extractor surpasses the state-of-the-art baseline on three imbalanced datasets; (2) the produced representations from our Emotion Extractor benefit the TTS model, and enable it to synthesize more expressive speech.
翻译:有效的语音情感表征在语音情感识别(SER)和情感文本转语音(TTS)任务中扮演关键角色。然而,与中性风格语音相比,情感语音样本的获取更加困难且成本更高,这导致大多数相关研究不幸忽略了一个问题:数据集非均衡。模型可能过度拟合占多数的中性类别,而无法生成鲁棒且有效的情感表征。本文提出一种情感提取器来解决该问题。我们采用增广方法训练模型,使其能够从非均衡数据集中提取有效且可泛化的情感表征。实证结果表明:(1)对于SER任务,所提出的情感提取器在三个非均衡数据集上超越了当前最优基线模型;(2)该情感提取器生成的表征有益于TTS模型,使其能够合成更具表现力的语音。