This paper investigates the optimal selection and fusion of feature encoders across multiple modalities and combines these in one neural network to improve sentiment detection. We compare different fusion methods and examine the impact of multi-loss training within the multi-modality fusion network, identifying surprisingly important findings relating to subnet performance. We have also found that integrating context significantly enhances model performance. Our best model achieves state-of-the-art performance for three datasets (CMU-MOSI, CMU-MOSEI and CH-SIMS). These results suggest a roadmap toward an optimized feature selection and fusion approach for enhancing sentiment detection in neural networks.
翻译:本文研究了跨多模态特征编码器的最优选择与融合方法,并将其整合于单一神经网络中以提升情感检测性能。我们比较了不同的融合策略,并探究了多模态融合网络中多损失训练的影响,发现了关于子网络性能的重要结论。研究还表明,上下文信息的整合能显著提升模型性能。我们提出的最优模型在三个数据集(CMU-MOSI、CMU-MOSEI和CH-SIMS)上取得了最先进的性能。这些结果为通过优化特征选择与融合方法来增强神经网络情感检测能力提供了技术路径。