Automatically recognising apparent emotions from face and voice is hard, in part because of various sources of uncertainty, including in the input data and the labels used in a machine learning framework. This paper introduces an uncertainty-aware audiovisual fusion approach that quantifies modality-wise uncertainty towards emotion prediction. To this end, we propose a novel fusion framework in which we first learn latent distributions over audiovisual temporal context vectors separately, and then constrain the variance vectors of unimodal latent distributions so that they represent the amount of information each modality provides w.r.t. emotion recognition. In particular, we impose Calibration and Ordinal Ranking constraints on the variance vectors of audiovisual latent distributions. When well-calibrated, modality-wise uncertainty scores indicate how much their corresponding predictions may differ from the ground truth labels. Well-ranked uncertainty scores allow the ordinal ranking of different frames across the modalities. To jointly impose both these constraints, we propose a softmax distributional matching loss. In both classification and regression settings, we compare our uncertainty-aware fusion model with standard model-agnostic fusion baselines. Our evaluation on two emotion recognition corpora, AVEC 2019 CES and IEMOCAP, shows that audiovisual emotion recognition can considerably benefit from well-calibrated and well-ranked latent uncertainty measures.
翻译:从面部和语音中自动识别外在情感是困难的,部分原因在于多种不确定性来源,包括输入数据以及机器学习框架中使用的标签。本文提出了一种不确定性感知的视听融合方法,可量化各模态对情感预测的不确定性。为此,我们设计了一种新型融合框架:首先分别学习视听时间上下文向量的潜在分布,然后约束单模态潜在分布的方差向量,使其表征各模态在情感识别中提供的信息量。具体而言,我们对视听潜在分布的方差向量施加了校准与序数排序约束。当达到良好校准时,模态层面的不确定性分数可表征其对应预测与真实标签的潜在差异程度;而良好的序数排序则允许跨模态对不同帧进行有序排列。为同时施加这两种约束,我们提出了基于Softmax的分布匹配损失函数。在分类与回归两种设定下,我们将所提出的不确定性感知融合模型与标准模型无关融合基线进行对比。在AVEC 2019 CES与IEMOCAP两个情感识别语料库上的评估表明,经过良好校准与排序的潜在不确定性度量可显著提升视听情感识别性能。