Pelican-Unified 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action

Yi Zhang,Yinda Chen,Che Liu,Zeyuan Ding,Jin Xu,Shilong Zou,Junwei Liao,Jiayu Hu,Xiancong Ren,Xiaopeng Zhang,Yechi Liu,Haoyuan Shi,Zecong Tang,Haosong Sun,Renwen Cui,Kuishu Wu,Wenhai Liu,Yang Xu,Yingji Zhang,Yidong Wang,Senkang Hu,Jinpeng Lu,Nga Teng Chan,Yechen Wu,Yong Dai,Jian Tang,Xiaozhu Ju

We present Pelican-Unified 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unified 1.0 uses a single VLM as a unified understanding module, mapping scenes, instructions, visual contexts, and action histories into a shared semantic space. The same VLM also serves as a unified reasoning module, autoregressively producing task-, action-, and future-oriented chains of thought in a single forward pass and projecting the final hidden state into a dense latent variable. A Unified Future Generator (UFG) then conditions on this latent variable and jointly generates future videos and future actions through two modality-specific output heads within the same denoising process. The language, video, and action losses are all backpropagated into the shared representation, enabling the model to jointly optimize understanding, reasoning, imagination, and action during training, rather than training three isolated expert systems. Experiments demonstrate that unification does not imply compromise. With a single checkpoint, Pelican-Unified 1.0 achieves strong performance across all three capabilities: 64.7 on eight VLM benchmarks, the best among comparable-scale models; 66.03 on WorldArena, ranking first; and 93.5 on RoboTwin, the second-best average among compared action methods. These results show that the unified paradigm succeeds in preserving specialist strength while bringing understanding, reasoning, imagination, and action into one model.

翻译：我们提出Pelican-Unified 1.0，这是首个依据统一化原则训练的具身基础模型。Pelican-Unified 1.0采用单一视觉语言模型（VLM）作为统一理解模块，将场景、指令、视觉上下文及动作历史映射至共享语义空间。该VLM同时作为统一推理模块，在一次前向传播中自回归地生成面向任务、动作及未来的思维链，并将最终隐状态投影为稠密潜变量。随后，统一未来生成器（UFG）以该潜变量为条件，在同一去噪过程中通过两个模态专用输出头联合生成未来视频与未来动作。语言、视频及动作损失均反向传播至共享表征，使模型在训练中能联合优化理解、推理、想象与行动，而非训练三个孤立的专家系统。实验表明，统一化并不意味着性能妥协。基于单一检查点，Pelican-Unified 1.0在全部三种能力上均取得强劲表现：在八项VLM基准测试中达64.7，为同规模模型最优；在WorldArena上以66.03分排名第一；在RoboTwin上达93.5，为所比较动作方法中第二佳均值。这些结果表明，统一范式成功保留了专家级性能，并将理解、推理、想象与行动融于单一模型。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

[ICML 2026] 看见的还是思考的？用奖励机制区分“看错”与“想错”：视觉语言模型奖励感知

专知会员服务

10+阅读 · 5月15日

【ICML2025】《引入推理于视觉：通过模型融合理解感知与推理》

专知会员服务

17+阅读 · 2025年5月12日

《面向无人机实时认知任务解决的视觉-语言-动作（VLA）模型与评估基准》

专知会员服务

42+阅读 · 2025年3月9日

【微软亚研】rStar-Math：小型大语言模型通过自我进化的深度思维掌握数学推理

专知会员服务

24+阅读 · 2025年1月13日