Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action

Yi Zhang,Yinda Chen,Che Liu,Zeyuan Ding,Jin Xu,Shilong Zou,Junwei Liao,Jiayu Hu,Xiancong Ren,Xiaopeng Zhang,Yechi Liu,Haoyuan Shi,Zecong Tang,Haosong Sun,Renwen Cui,Kuishu Wu,Wenhai Liu,Yang Xu,Yingji Zhang,Yidong Wang,Senkang Hu,Jinpeng Lu,Nga Teng Chan,Yechen Wu,Zeting Liu,Xianzhou Hou,Yong Dai,Jian Tang,Xiaozhu Ju

We present Pelican-Unify 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unify 1.0 uses a single VLM as a unified understanding module, mapping scenes, instructions, visual contexts, and action histories into a shared semantic space. The same VLM also serves as a unified reasoning module, autoregressively producing task-, action-, and future-oriented chains of thought in a single forward pass and projecting the final hidden state into a dense latent variable. A Unified Future Generator (UFG) then conditions on this latent variable and jointly generates future videos and future actions through two modality-specific output heads within the same denoising process. The language, video, and action losses are all backpropagated into the shared representation, enabling the model to jointly optimize understanding, reasoning, imagination, and action during training, rather than training three isolated expert systems. Experiments demonstrate that unification does not imply compromise. With a single checkpoint, Pelican-Unify 1.0 achieves strong performance across all three capabilities: 64.7 on eight VLM benchmarks, the best among comparable-scale models; 66.03 on WorldArena, ranking first; and 93.5 on RoboTwin, the second-best average among compared action methods. These results show that the unified paradigm succeeds in preserving specialist strength while bringing understanding, reasoning, imagination, and action into one model.

翻译：我们提出Pelican-Unify 1.0，这是首个遵循统一原则训练的具身基础模型。Pelican-Unify 1.0采用单一VLM作为统一理解模块，将场景、指令、视觉上下文和动作历史映射到共享语义空间。同一VLM也作为统一推理模块，在单次前向传递中自回归生成面向任务、动作和未来的思维链，并将最终隐藏状态投影为稠密潜变量。随后，统一未来生成器（UFG）以该潜变量为条件，在相同的去噪过程中通过两个模态专用输出头联合生成未来视频和未来动作。语言、视频和动作损失均反向传播至共享表征，使模型能够在训练时联合优化理解、推理、想象和行动，而非训练三个孤立的专家系统。实验表明，统一并不意味着妥协。通过单一检查点，Pelican-Unify 1.0在三种能力上均展现出强劲性能：在八项VLM基准测试中得分64.7，为可比规模模型中的最优；在WorldArena上得分66.03，排名第一；在RoboTwin上得分93.5，在比较的动作方法中平均分排名第二。这些结果表明，统一范式在保持专家优势的同时，成功将理解、推理、想象和行动整合至单一模型中。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

[ICML 2026] 看见的还是思考的？用奖励机制区分“看错”与“想错”：视觉语言模型奖励感知

专知会员服务

10+阅读 · 5月15日

【微软亚研】rStar-Math：小型大语言模型通过自我进化的深度思维掌握数学推理

专知会员服务

24+阅读 · 2025年1月13日

VILA-U：一个融合视觉理解与生成的统一基础模型

专知会员服务

21+阅读 · 2024年9月9日