Multimodal foundation models have demonstrated impressive capabilities across diverse tasks. However, their potential as plug-and-play solutions for missing modality reconstruction remains underexplored. To bridge this gap, we identify and formalize three potential paradigms for missing modality reconstruction, and perform a comprehensive evaluation across these paradigms, covering 42 model variants in terms of reconstruction accuracy and adaptability to downstream tasks. Our analysis reveals that current foundation models often fall short in two critical aspects: (i) fine-grained semantic extraction from the available modalities, and (ii) robust validation of generated modalities. These limitations lead to suboptimal and, at times, misaligned generations. To address these challenges, we propose an agentic framework tailored for missing modality reconstruction. This framework dynamically formulates modality-aware mining strategies based on the input context, facilitating the extraction of richer and more discriminative semantic features. In addition, we introduce a self-refinement mechanism, which iteratively verifies and enhances the quality of generated modalities through internal feedback. Experimental results show that our method reduces FID for missing image reconstruction by at least 14\% and MER for missing text reconstruction by at least 10\% compared to baselines. Code are released at: https://github.com/Guanzhou-Ke/AFM2.
翻译:多模态基础模型已在各类任务中展现出令人印象深刻的能力。然而,它们作为即插即用解决方案用于缺失模态重建的潜力仍未得到充分探索。为弥合这一差距,我们识别并形式化了三种潜在的缺失模态重建范式,并针对这些范式进行了全面评估,涵盖42种模型变体,在重建精度和下游任务适应性方面进行了分析。我们的分析表明,当前的基础模型在两个关键方面往往表现欠佳:(i) 从可用模态中提取细粒度语义,以及 (ii) 对生成模态的稳健验证。这些限制导致了次优甚至有时不协调的生成结果。为应对这些挑战,我们提出了一种专为缺失模态重建设计的智能体框架。该框架基于输入上下文动态制定模态感知挖掘策略,有助于提取更丰富、更具判别性的语义特征。此外,我们引入了一种自改进机制,通过内部反馈迭代验证并提升生成模态的质量。实验结果显示,与基线方法相比,我们的方法在缺失图像重建上至少将FID降低了14%,在缺失文本重建上至少将MER降低了10%。代码已发布于:https://github.com/Guanzhou-Ke/AFM2。