Affective computing increasingly relies on deep learning to represent emotions, yet latent spaces often remain opaque, high-dimensional black boxes. This paper investigates whether Transformers' embeddings recover the geometric regularities of Russell's circumplex model. We unify two complementary experiments testing the hypothesis that, after training models on text and speech, their resulting latent spaces encode a topology consistent with valence-arousal and reproduce human-like neighborhood relations. Specifically, we evaluate deep representations extracted from Transformer-based text (RoBERTa) and speech (wav2vec 2.0) encoders, along with a multimodal Transformer fusion architecture, across naturalistic datasets like MSP-Podcast and controlled LLM-generated stimuli. Our analysis reveals that multimodal fusion of text and audio yields perfect topological alignment with Russell's primary emotion ordering. Furthermore, in a zero-shot setting using generic text embeddings, projected fine-grained emotion terms fall close to their established human-mapped coordinates. Our contribution is a novel, data-driven framework for validating emotion models, demonstrating that Russell's circumplex structure is intrinsically encoded in the embeddings of these modalities rather than being solely an artifact of human labeling, thereby bridging the gap between psychological theory and representation learning.
翻译:情感计算日益依赖深度学习来表征情感,然而潜在空间往往仍是黑盒,且维度较高。本文探究Transformer的嵌入是否能恢复罗素环状模型的几何规律。我们统一了两个互补实验,以验证以下假设:在基于文本和语音训练模型后,其潜在空间编码的拓扑结构与效价-唤醒维度一致,并能复现类似人类的邻近关系。具体而言,我们评估了从基于Transformer的文本编码器(RoBERTa)和语音编码器(wav2vec 2.0)以及多模态Transformer融合架构中提取的深度表征,数据集涵盖MSP-Podcast等自然语料库和LLM生成的受控刺激。分析表明,文本与音频的多模态融合在拓扑上与罗素的主要情感排序完美对齐。此外,在零样本设置下,使用通用文本嵌入时,投影后的细粒度情感术语紧邻其已有的人类映射坐标。我们的贡献在于提出了一种新颖的数据驱动情感模型验证框架,证明罗素环状结构内在地编码于这些模态的嵌入中,而非仅仅是人类标注的产物,从而弥合了心理学理论与表征学习之间的鸿沟。