We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.
翻译:我们提出了Intern-S1-Pro,首个参数量达一万亿的科学多模态基础模型。通过扩展至这一前所未有的规模,该模型在通用与科学领域均实现了全面性能提升。除更强的推理与图像-文本理解能力外,其智能体能力得到增强,同时科学专业知识大幅拓展,可掌握化学、材料、生命科学与地球科学等关键科学领域超过100项专业任务。实现这一超大规模得益于XTuner与LMDeploy的强大基础设施支持,前者可在万亿参数级别实现高度高效的强化学习训练,后者则确保训练与推理间严格的精度一致性。通过无缝整合这些突破,Intern-S1-Pro进一步强化了通用智能与专业智能的融合,以“可专业化通才”之姿展现效能:在通用能力上跻身开源模型第一梯队,在科学专业任务深度上超越闭源模型。