Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe that autoregressive normalizing flows are autoregressive Transformers--sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs--making them the most natural paradigm for true unified multimodal generation. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a pretrained VLM stream with a TarFlow stream via residual skip connections, both operating under the same causal mask. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 enables cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.


翻译:深度生成模型在文本与视觉领域发展迅速,催生了能够理解、推理并生成交错的文本-图像序列的统一多模态系统。现有方法多将自回归语言建模与基于扩散的图像生成器结合,但因果文本生成与迭代式视觉去噪之间存在结构性不匹配。我们注意到,自回归归一化流本质上就是自回归Transformer——与LLM共享相同的因果掩码、KV缓存机制及从左到右的结构——使其成为实现真正统一多模态生成的最自然范式。我们提出了STARFlow2,它基于Pretzel架构,通过残差跳跃连接将预训练的视觉语言模型(VLM)流与TarFlow流垂直交错,二者均在同一因果掩码下运行。结合深度-浅层流设计与统一FAE潜在空间,STARFlow2实现了缓存友好的交错式生成,其中文本与视觉输出可直接进入KV缓存而无需重新编码。实验表明,该模型在图像生成与多模态理解基准上均展现出强劲性能,验证了自回归流作为统一多模态建模可行基础的有效性。

0
下载
关闭预览

相关内容

144页ppt《扩散模型》,Google DeepMind Sander Dieleman
专知会员服务
51+阅读 · 2025年11月21日
多模态大型语言模型:综述
专知会员服务
47+阅读 · 2025年6月14日
统一的多模态理解与生成模型:进展、挑战与机遇
专知会员服务
34+阅读 · 2025年5月6日
《多模态大型语言模型进化》最新综述
专知会员服务
105+阅读 · 2024年2月23日
使用多模态语言模型生成图像
专知会员服务
32+阅读 · 2023年8月23日
Meta-Transformer:多模态学习的统一框架
专知会员服务
59+阅读 · 2023年7月21日
专家报告|深度学习+图像多模态融合
中国图象图形学报
12+阅读 · 2019年10月23日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
8+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
10+阅读 · 8月1日
相关VIP内容
144页ppt《扩散模型》,Google DeepMind Sander Dieleman
专知会员服务
51+阅读 · 2025年11月21日
多模态大型语言模型:综述
专知会员服务
47+阅读 · 2025年6月14日
统一的多模态理解与生成模型:进展、挑战与机遇
专知会员服务
34+阅读 · 2025年5月6日
《多模态大型语言模型进化》最新综述
专知会员服务
105+阅读 · 2024年2月23日
使用多模态语言模型生成图像
专知会员服务
32+阅读 · 2023年8月23日
Meta-Transformer:多模态学习的统一框架
专知会员服务
59+阅读 · 2023年7月21日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员