This open access book provides a comprehensive overview of the state of the art in research and applications of Foundation Models and is intended for readers familiar with basic Natural Language Processing (NLP) concepts. Over the recent years, a revolutionary new paradigm has been developed for training models for NLP. These models are first pre-trained on large collections of text documents to acquire general syntactic knowledge and semantic information. Then, they are fine-tuned for specific tasks, which they can often solve with superhuman accuracy. When the models are large enough, they can be instructed by prompts to solve new tasks without any fine-tuning. Moreover, they can be applied to a wide range of different media and problem domains, ranging from image and video processing to robot control learning. Because they provide a blueprint for solving many tasks in artificial intelligence, they have been called Foundation Models. After a brief introduction to basic NLP models the main pre-trained language models BERT, GPT and sequence-to-sequence transformer are described, as well as the concepts of self-attention and context-sensitive embedding. Then, different approaches to improving these models are discussed, such as expanding the pre-training criteria, increasing the length of input texts, or including extra knowledge. An overview of the best-performing models for about twenty application areas is then presented, e.g., question answering, translation, story generation, dialog systems, generating images from text, etc. For each application area, the strengths and weaknesses of current models are discussed, and an outlook on further developments is given. In addition, links are provided to freely available program code. A concluding chapter summarizes the economic opportunities, mitigation of risks, and potential developments of AI.
翻译:这本开放获取图书全面概述了基础模型在研究和应用方面的最新进展,旨在为熟悉自然语言处理基础概念的读者编写。近年来,一种用于训练NLP模型的革命性新范式得以发展。这些模型首先在大量文本文档集合上进行预训练,以获取通用句法知识和语义信息。随后,它们针对特定任务进行微调,通常能以超越人类的准确率解决这些问题。当模型规模足够大时,它们可以通过提示被引导解决新任务,而无需任何微调。此外,它们还可广泛应用于从图像与视频处理到机器人控制学习等不同媒体和问题领域。由于这些模型为解决人工智能中的众多任务提供了蓝图,因此被称为基础模型。在简要介绍基础NLP模型之后,本书描述了主要的预训练语言模型BERT、GPT和序列到序列Transformer,以及自注意力和上下文敏感嵌入的概念。接着,讨论了改进这些模型的不同方法,例如扩展预训练标准、增加输入文本长度或引入额外知识。随后,概述了约二十个应用领域中性能最佳的模型,例如问答、翻译、故事生成、对话系统、从文本生成图像等。针对每个应用领域,讨论了当前模型的优缺点,并展望了进一步的发展方向。此外,还提供了免费可用的程序代码链接。最后一章总结了人工智能的经济机遇、风险缓解措施及潜在发展。