Human emotion is expressed in many communication modalities and media formats and so their computational study is equally diversified into natural language processing, audio signal analysis, computer vision, etc. Similarly, the large variety of representation formats used in previous research to describe emotions (polarity scales, basic emotion categories, dimensional approaches, appraisal theory, etc.) have led to an ever proliferating diversity of datasets, predictive models, and software tools for emotion analysis. Because of these two distinct types of heterogeneity, at the expressional and representational level, there is a dire need to unify previous work on increasingly diverging data and label types. This article presents such a unifying computational model. We propose a training procedure that learns a shared latent representation for emotions, so-called emotion embeddings, independent of different natural languages, communication modalities, media or representation label formats, and even disparate model architectures. Experiments on a wide range of heterogeneous affective datasets indicate that this approach yields the desired interoperability for the sake of reusability, interpretability and flexibility, without penalizing prediction quality. Code and data are archived under https://doi.org/10.5281/zenodo.7405327 .
翻译:人类情感通过多种交流模态和媒体形式表达,因此在计算方法研究中也广泛涉及自然语言处理、音频信号分析、计算机视觉等不同领域。同样,以往研究中用于描述情感的各种表征格式(如极性量表、基本情感类别、维度方法、评价理论等)导致了情感分析领域数据集、预测模型及软件工具的持续多样化。由于表达层面和表征层面存在这两种不同类型的异质性,当前亟需统一处理日益分化的数据和标签类型。本文提出了一种统一的计算模型:我们设计了一种训练流程,用于学习情感的共享潜在表征(称为情感嵌入),该嵌入独立于不同自然语言、交流模态、媒体或表征标签格式,甚至不同模型架构。在广泛的异质情感数据集上的实验表明,该方法能够在不牺牲预测质量的前提下,实现理想的互操作性,从而提升可重用性、可解释性和灵活性。代码和数据存档于 https://doi.org/10.5281/zenodo.7405327 。