In this paper, we present a multimodal \textit{and} dynamical VAE (MDVAE) applied to unsupervised audio-visual speech representation learning. The latent space is structured to dissociate the latent dynamical factors that are shared between the modalities from those that are specific to each modality. A static latent variable is also introduced to encode the information that is constant over time within an audiovisual speech sequence. The model is trained in an unsupervised manner on an audiovisual emotional speech dataset, in two stages. In the first stage, a vector quantized VAE (VQ-VAE) is learned independently for each modality, without temporal modeling. The second stage consists in learning the MDVAE model on the intermediate representation of the VQ-VAEs before quantization. The disentanglement between static versus dynamical and modality-specific versus modality-common information occurs during this second training stage. Extensive experiments are conducted to investigate how audiovisual speech latent factors are encoded in the latent space of MDVAE. These experiments include manipulating audiovisual speech, audiovisual facial image denoising, and audiovisual speech emotion recognition. The results show that MDVAE effectively combines the audio and visual information in its latent space. They also show that the learned static representation of audiovisual speech can be used for emotion recognition with few labeled data, and with better accuracy compared with unimodal baselines and a state-of-the-art supervised model based on an audiovisual transformer architecture.
翻译:本文提出了一种多模态动态变分自编码器(MDVAE),并将其应用于无监督的视听语音表征学习。其潜在空间被结构化为能够分离模态间共享的动态潜在因子与各模态特有动态潜在因子。同时引入静态潜在变量以编码视听语音序列中随时间恒定不变的信息。该模型以无监督方式在视听情感语音数据集上分两阶段训练:第一阶段为各模态独立学习向量量化变分自编码器(VQ-VAE),不包含时间建模;第二阶段则在VQ-VAE的量化前中间表征上学习MDVAE模型。静态与动态、模态特有与模态共享信息的解耦发生在第二阶段训练过程中。通过大量实验探究视听语音潜在因子在MDVAE潜在空间中的编码方式,实验涵盖视听语音操控、视听人脸图像去噪以及视听语音情感识别。结果表明,MDVAE在其潜在空间中有效融合了音频与视觉信息,且学习到的视听语音静态表征可用于小样本情感识别,其准确率优于单模态基线模型及基于视听Transformer架构的现有最佳监督模型。