Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. Current research is fragmented, with studies varying significantly across datasets, modality configurations, and experimental settings. Furthermore, existing benchmarks focus predominantly on action recognition, often neglecting critical real-world challenges such as input corruptions, missing modalities, and model trustworthiness. This lack of standardization obscures a reliable assessment of the field's advancement. To address this issue, we introduce MMDG-Bench, the first unified and comprehensive benchmark for MMDG, which standardizes evaluation across six datasets spanning three diverse tasks: action recognition, mechanical fault diagnosis, and sentiment analysis. MMDG-Bench encompasses six modality combinations, nine representative methods, and multiple evaluation settings. Beyond standard accuracy, it systematically assesses corruption robustness, missing-modality generalization, misclassification detection, and out-of-distribution detection. With 7, 402 neural networks trained in total across 95 unique cross-domain tasks, MMDG-Bench yields five key findings: (1) under fair comparisons, recent specialized MMDG methods offer only marginal improvements over ERM baseline; (2) no single method consistently outperforms others across datasets or modality combinations; (3) a substantial gap to upper-bound performance persists, indicating that MMDG remains far from solved; (4) trimodal fusion does not consistently outperform the strongest bimodal configurations; and (5) all evaluated methods exhibit significant degradation under corruption and missing-modality scenarios, with some methods further compromising model trustworthiness.


翻译:尽管多模态领域泛化在提升模型鲁棒性方面日益受到关注,但现有性能提升究竟是算法层面的实质性进展,还是评估协议不统一造成的人为假象,这一问题尚不明确。当前研究呈现碎片化特征,不同工作在数据集、模态配置和实验设置上存在显著差异。此外,现有基准主要聚焦动作识别任务,往往忽略输入损坏、模态缺失和模型可信度等关键现实挑战。这种标准化缺失阻碍了对领域发展水平的可靠评估。为解决该问题,我们提出MMDG-Bench——首个统一且全面的多模态领域泛化基准。该基准在涵盖动作识别、机械故障诊断和情感分析三项不同任务的六个数据集上实现标准化评估。MMDG-Bench包含六种模态组合、九种代表性方法及多种评估设置。除标准准确率外,本基准系统性地评估了损坏鲁棒性、缺失模态泛化能力、误分类检测及分布外检测性能。通过训练总计7,402个神经网络(覆盖95个独特的跨领域任务),MMDG-Bench得出五项关键发现:(1)在公平比较条件下,近期专用多模态领域泛化方法相比ERM基线仅带来边际提升;(2)没有任何单一方法能在不同数据集或模态组合中持续表现最优;(3)与最优性能上限仍存在显著差距,表明多模态领域泛化问题远未解决;(4)三模态融合并非始终优于最强双模态配置;(5)所有评估方法在损坏和缺失模态场景下均出现显著性能退化,部分方法进一步损害了模型可信度。

0
下载
关闭预览

相关内容

【CMU博士论文】深度学习中泛化的量化、理解与改进
专知会员服务
21+阅读 · 2025年10月11日
深度学习中泛化的量化、理解与改进
专知会员服务
17+阅读 · 2025年9月13日
探究模型能力与应用的进展和边界
专知会员服务
26+阅读 · 2025年8月27日
《多模态适应与泛化》进展综述:从传统方法到基础模型
多模态大规模语言模型基准的综述
专知会员服务
41+阅读 · 2024年8月25日
多智能体深度强化学习研究进展
专知会员服务
76+阅读 · 2024年7月17日
多模态大语言模型研究进展!
专知会员服务
43+阅读 · 2024年7月15日
专知会员服务
53+阅读 · 2021年8月13日
基于深度学习的多标签生成研究进展
专知会员服务
148+阅读 · 2020年4月25日
「基于通信的多智能体强化学习」 进展综述
多模态视觉语言表征学习研究综述
专知
27+阅读 · 2020年12月3日
深度多模态表示学习综述论文,22页pdf
专知
33+阅读 · 2020年6月21日
专访俞栋:多模态是迈向通用人工智能的重要方向
AI科技评论
27+阅读 · 2019年9月9日
这可能是「多模态机器学习」最通俗易懂的介绍
计算机视觉life
113+阅读 · 2018年12月20日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Arxiv
14+阅读 · 2024年5月28日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
1+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
【CMU博士论文】深度学习中泛化的量化、理解与改进
专知会员服务
21+阅读 · 2025年10月11日
深度学习中泛化的量化、理解与改进
专知会员服务
17+阅读 · 2025年9月13日
探究模型能力与应用的进展和边界
专知会员服务
26+阅读 · 2025年8月27日
《多模态适应与泛化》进展综述:从传统方法到基础模型
多模态大规模语言模型基准的综述
专知会员服务
41+阅读 · 2024年8月25日
多智能体深度强化学习研究进展
专知会员服务
76+阅读 · 2024年7月17日
多模态大语言模型研究进展!
专知会员服务
43+阅读 · 2024年7月15日
专知会员服务
53+阅读 · 2021年8月13日
基于深度学习的多标签生成研究进展
专知会员服务
148+阅读 · 2020年4月25日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员