Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where generation facilitates understanding. To this end, we introduce UniG2U-Bench, a comprehensive benchmark categorizing generation-to-understanding (G2U) evaluation into 7 regimes and 30 subtasks, requiring varying degrees of implicit or explicit visual transformations. Extensive evaluation of over 30 models reveals three core findings: 1) Unified models generally underperform their base Vision-Language Models (VLMs), and Generate-then-Answer (GtA) inference typically degrades performance relative to direct inference. 2) Consistent enhancements emerge in spatial intelligence, visual illusions, or multi-round reasoning subtasks, where enhanced spatial and shape perception, as well as multi-step intermediate image states, prove beneficial. 3) Tasks with similar reasoning structures and models sharing architectures exhibit correlated behaviors, suggesting that generation-understanding coupling induces class-consistent inductive biases over tasks, pretraining data, and model architectures. These findings highlight the necessity for more diverse training data and novel paradigms to fully unlock the potential of unified multimodal modeling.


翻译:统一多模态模型近期展现出强大的生成能力,然而生成是否以及何时能够提升理解能力,目前尚不明确。现有基准测试缺乏对生成促进理解的具体任务进行系统性探索。为此,我们提出了UniG2U-Bench,这是一个综合性基准测试,将生成到理解(G2U)的评估划分为7种范式与30个子任务,这些任务需要不同程度的隐式或显式视觉转换。通过对超过30个模型进行广泛评估,我们得出了三个核心发现:1)统一模型通常表现逊于其基础视觉语言模型(VLMs),且“生成后回答”(GtA)推理相较于直接推理通常会降低性能。2)在空间智能、视觉错觉或多轮推理子任务中,模型表现出一致的提升,其中增强的空间与形状感知能力以及多步中间图像状态被证明是有益的。3)具有相似推理结构的任务以及共享架构的模型表现出相关性行为,这表明生成与理解的耦合在任务、预训练数据和模型架构上诱导了类别一致的归纳偏置。这些发现凸显了需要更多样化的训练数据和新颖的范式,以充分释放统一多模态建模的潜力。

0
下载
关闭预览

相关内容

CVPR 2026教程:统一多模态模型走向收敛之路
专知会员服务
20+阅读 · 6月8日
【博士论文】弥合多模态基础模型与世界模型之间的鸿沟
统一的多模态理解与生成模型:进展、挑战与机遇
专知会员服务
34+阅读 · 2025年5月6日
对比预训练和多模态生成式人工智能的统计理论
专知会员服务
23+阅读 · 2025年1月12日
统一的多模态文字理解与生成大模型
专知会员服务
30+阅读 · 2024年10月11日
Meta-Transformer:多模态学习的统一框架
专知会员服务
59+阅读 · 2023年7月21日
专访俞栋:多模态是迈向通用人工智能的重要方向
AI科技评论
27+阅读 · 2019年9月9日
这可能是「多模态机器学习」最通俗易懂的介绍
计算机视觉life
113+阅读 · 2018年12月20日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
6+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
7+阅读 · 7月19日
战力倍增器:自主武器系统与乌克兰及加沙冲突
人工智能赋能战场情报:提速决策进程
专知会员服务
5+阅读 · 7月17日
《拥抱新兴技术:面向未来军官的教育革新》
专知会员服务
8+阅读 · 7月17日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员