Human emotion understanding is pivotal in making conversational technology mainstream. We view speech emotion understanding as a perception task which is a more realistic setting. With varying contexts (languages, demographics, etc.) different share of people perceive the same speech segment as a non-unanimous emotion. As part of the ACM Multimedia 2023 Computational Paralinguistics ChallengE (ComParE) in the EMotion Share track, we leverage their rich dataset of multilingual speakers and multi-label regression target of 'emotion share' or perception of that emotion. We demonstrate that the training scheme of different foundation models dictates their effectiveness for tasks beyond speech recognition, especially for non-semantic speech tasks like emotion understanding. This is a very complex task due to multilingual speakers, variability in the target labels, and inherent imbalance in the regression dataset. Our results show that HuBERT-Large with a self-attention-based light-weight sequence model provides 4.6% improvement over the reported baseline.
翻译:人类情感理解是推动对话式技术普及的关键。我们将语音情感理解视为感知任务,这更贴近实际场景。在不同语境(如语言、人口统计特征等)下,不同人群对同一语音片段的感知情感存在差异性。在ACM Multimedia 2023计算副语言学挑战赛(ComParE)的情感共享赛道中,我们利用其丰富的多语种说话人数据集以及多标签回归目标——即"情感共享度"或对该情感的感知强度。研究表明,不同基础模型的训练策略决定了它们在语音识别之外任务上的有效性,尤其是针对情感理解这类非语义语音任务。由于多语种说话人、目标标签的变异性以及回归数据集固有的不平衡性,该任务极具挑战性。我们的实验结果显示,基于自注意力机制的轻量级序列模型HuBERT-Large相比基线方案实现了4.6%的性能提升。