Speech emotion recognition (SER) systems can exhibit gender-related performance disparities, but how such bias manifests in multilingual speech LLMs across languages and modalities is unclear. We introduce a novel multilingual, multimodal benchmark built on MELD-ST, spanning English, Japanese, and German, to quantify language-specific SER performance and gender gaps. We find bias is strongly language-dependent, and multimodal fusion does not reliably improve fairness. To address these, we propose ERM-MinMaxGAP, a fairness-informed training objective, which augments empirical risk minimization (ERM) with a proposed adaptive fairness weight mechanism and a novel MinMaxGAP regularizer on the maximum male-female loss gap within each language and modality. Building upon the Qwen2-Audio backbone, our ERM-MinMaxGAP approach improves multilingual SER performance by 5.5% and 5.0% while reducing the overall gender bias gap by 0.1% and 1.4% in the unimodal and multimodal settings, respectively.
翻译:语音情感识别系统可能表现出与性别相关的性能差异,但这种偏见如何在跨语言和跨模态的多语言语音大模型中体现尚不明确。我们基于MELD-ST构建了一个新颖的多语言多模态基准数据集,涵盖英语、日语和德语,用于量化特定语言的情感识别性能与性别差距。研究发现,偏见强烈依赖于语言,且多模态融合并不能可靠地提升公平性。为解决这些问题,我们提出了ERM-MinMaxGAP——一种面向公平性的训练目标,它通过引入自适应公平性权重机制和新型MinMaxGAP正则化项(作用于每种语言和模态内的最大性别损失差距)对经验风险最小化进行增强。基于Qwen2-Audio骨干网络,我们的ERM-MinMaxGAP方法在单模态和多模态设置下,分别将多语言情感识别性能提升了5.5%和5.0%,同时将整体性别偏见差距降低了0.1%和1.4%。