Large Language Models (LLMs) have become ubiquitous across various domains, transforming the way we interact with information and conduct research. However, most high-performing LLMs remain confined behind proprietary walls, hindering scientific progress. Most open-source LLMs, on the other hand, are limited in their ability to support longer sequence lengths, which is a key requirement for many tasks that require inference over an input context. To address this, we have trained XGen, a series of 7B parameter models on up to 8K sequence length for up to 1.5T tokens. We have also finetuned the XGen models on public-domain instructional data, creating their instruction-tuned counterparts (XGen-Inst). We open-source our models for both research advancements and commercial applications. Our evaluation on standard benchmarks shows that XGen models achieve comparable or better results when compared with state-of-the-art open-source LLMs. Our targeted evaluation on long sequence modeling tasks shows the benefits of our 8K-sequence models over 2K-sequence open-source LLMs.
翻译:大型语言模型(LLMs)已在各个领域变得无处不在,彻底改变了我们与信息交互及开展研究的方式。然而,大多数高性能大语言模型仍被限制在专有壁垒之内,阻碍了科学进步。另一方面,大多数开源大语言模型在支持较长序列长度方面能力有限,而这对许多需要基于输入上下文进行推理的任务而言是关键要求。为解决这一问题,我们训练了XGen系列模型(拥有7B参数),其序列长度达8K,训练数据规模达1.5T词元。我们还利用公共领域指令数据对XGen模型进行了微调,生成了对应的指令调优版本(XGen-Inst)。我们将这些模型开源,以同时促进研究进展和商业应用。在标准基准测试上的评估表明,XGen模型与最先进的开源大语言模型相比表现相当或更优。针对长序列建模任务的专项评估显示,我们的8K序列模型相较于2K序列的开源大语言模型具有显著优势。