Most music emotion recognition approaches perform classification or regression that estimates a general emotional category from a distribution of music samples, but without considering emotional variations (e.g., happiness can be further categorised into much, moderate or little happiness). We propose an embedding-based music emotion recognition approach that associates music samples with emotions in a common embedding space by considering both general emotional categories and fine-grained discrimination within each category. Since the association of music samples with emotions is uncertain due to subjective human perceptions, we compute composite loss-based embeddings obtained to maximise two statistical characteristics, one being the correlation between music samples and emotions based on canonical correlation analysis, and the other being a probabilistic similarity between a music sample and an emotion with KL-divergence. The experiments on two benchmark datasets demonstrate the effectiveness of our embedding-based approach, the composite loss and learned acoustic features. In addition, detailed analysis shows that our approach can accomplish robust bidirectional music emotion recognition that not only identifies music samples matching with a specific emotion but also detects emotions expressed in a certain music sample.
翻译:现有的大多数音乐情感识别方法通过分类或回归从音乐样本分布中估计大致的情感类别,但未能考虑情感内部的细微差异(例如,快乐可进一步分为非常快乐、中等快乐或略带快乐)。本文提出一种基于嵌入的音乐情感识别方法,通过兼顾大致情感类别与类别内的精细区分,将音乐样本与情感关联至统一的嵌入空间。由于人为感知的主观性导致音乐样本与情感的关联存在不确定性,我们计算基于复合损失的嵌入向量,其优化目标包含两种统计特性:一是基于典型相关分析的音乐样本与情感之间的相关性,二是通过KL散度衡量的音乐样本与情感之间的概率相似度。在两个基准数据集上的实验验证了所提基于嵌入的方法、复合损失函数以及学习到的声学特征的有效性。此外,详细分析表明,该方法能够实现鲁棒的双向音乐情感识别——既能为特定情感匹配对应的音乐样本,也能检测特定音乐样本所表达的情感。