Large Language Models (LLMs) have drawn a lot of attention due to their strong performance on a wide range of natural language tasks, since the release of ChatGPT in November 2022. LLMs' ability of general-purpose language understanding and generation is acquired by training billions of model's parameters on massive amounts of text data, as predicted by scaling laws \cite{kaplan2020scaling,hoffmann2022training}. The research area of LLMs, while very recent, is evolving rapidly in many different ways. In this paper, we review some of the most prominent LLMs, including three popular LLM families (GPT, LLaMA, PaLM), and discuss their characteristics, contributions and limitations. We also give an overview of techniques developed to build, and augment LLMs. We then survey popular datasets prepared for LLM training, fine-tuning, and evaluation, review widely used LLM evaluation metrics, and compare the performance of several popular LLMs on a set of representative benchmarks. Finally, we conclude the paper by discussing open challenges and future research directions.
翻译:大型语言模型(LLMs)自2022年11月ChatGPT发布以来,凭借其在各类自然语言任务中的卓越表现吸引了广泛关注。正如标度定律\cite{kaplan2020scaling,hoffmann2022training}所预测的那样,LLMs通过在海量文本数据上对数十亿参数进行训练,获得了通用语言理解与生成能力。尽管该研究领域尚属新兴,但其发展路径多元且演进迅速。本文系统梳理了最具代表性的LLMs,涵盖三大主流模型族(GPT、LLaMA、PaLM),并探讨了它们的特性、贡献与局限性。我们概述了构建与增强LLMs的关键技术,系统整理了用于LLM训练、微调与评估的主流数据集,评述了广泛采用的评估指标,并在代表性基准上对比了若干知名LLMs的性能。最后,本文总结了当前面临的挑战与未来研究方向。