Multimodal foundation models have demonstrated impressive capabilities across diverse tasks. However, their potential as plug-and-play solutions for missing modality reconstruction remains underexplored. To bridge this gap, we identify and formalize three potential paradigms for missing modality reconstruction, and perform a comprehensive evaluation across these paradigms, covering 42 model variants in terms of reconstruction accuracy and adaptability to downstream tasks. Our analysis reveals that current foundation models often fall short in two critical aspects: (i) fine-grained semantic extraction from the available modalities, and (ii) robust validation of generated modalities. These limitations lead to suboptimal and, at times, misaligned generations. To address these challenges, we propose an agentic framework tailored for missing modality reconstruction. This framework dynamically formulates modality-aware mining strategies based on the input context, facilitating the extraction of richer and more discriminative semantic features. In addition, we introduce a self-refinement mechanism, which iteratively verifies and enhances the quality of generated modalities through internal feedback. Experimental results show that our method reduces FID for missing image reconstruction by at least 14\% and MER for missing text reconstruction by at least 10\% compared to baselines. Code are released at: https://github.com/Guanzhou-Ke/AFM2.


翻译:多模态基础模型已在各类任务中展现出令人印象深刻的能力。然而,它们作为即插即用解决方案用于缺失模态重建的潜力仍未得到充分探索。为弥合这一差距,我们识别并形式化了三种潜在的缺失模态重建范式,并针对这些范式进行了全面评估,涵盖42种模型变体,在重建精度和下游任务适应性方面进行了分析。我们的分析表明,当前的基础模型在两个关键方面往往表现欠佳:(i) 从可用模态中提取细粒度语义,以及 (ii) 对生成模态的稳健验证。这些限制导致了次优甚至有时不协调的生成结果。为应对这些挑战,我们提出了一种专为缺失模态重建设计的智能体框架。该框架基于输入上下文动态制定模态感知挖掘策略,有助于提取更丰富、更具判别性的语义特征。此外,我们引入了一种自改进机制,通过内部反馈迭代验证并提升生成模态的质量。实验结果显示,与基线方法相比,我们的方法在缺失图像重建上至少将FID降低了14%,在缺失文本重建上至少将MER降低了10%。代码已发布于:https://github.com/Guanzhou-Ke/AFM2。

0
下载
关闭预览

相关内容

【博士论文】弥合多模态基础模型与世界模型之间的鸿沟
【CVPR2025】知识桥接器:走向无训练的缺失模态补全
专知会员服务
14+阅读 · 2025年2月28日
遥感基础模型发展综述与未来设想
专知会员服务
20+阅读 · 2024年8月13日
深度多模态表示学习综述论文,22页pdf
专知
33+阅读 · 2020年6月21日
专访俞栋:多模态是迈向通用人工智能的重要方向
AI科技评论
27+阅读 · 2019年9月9日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
如何理解模型的过拟合与欠拟合,以及如何解决?
七月在线实验室
12+阅读 · 2019年4月23日
用模型不确定性理解模型
论智
11+阅读 · 2018年9月5日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
Arxiv
14+阅读 · 2024年5月28日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
3+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
5+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
8+阅读 · 7月19日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员