We introduce PaLM 2, a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM. PaLM 2 is a Transformer-based model trained using a mixture of objectives. Through extensive evaluations on English and multilingual language, and reasoning tasks, we demonstrate that PaLM 2 has significantly improved quality on downstream tasks across different model sizes, while simultaneously exhibiting faster and more efficient inference compared to PaLM. This improved efficiency enables broader deployment while also allowing the model to respond faster, for a more natural pace of interaction. PaLM 2 demonstrates robust reasoning capabilities exemplified by large improvements over PaLM on BIG-Bench and other reasoning tasks. PaLM 2 exhibits stable performance on a suite of responsible AI evaluations, and enables inference-time control over toxicity without additional overhead or impact on other capabilities. Overall, PaLM 2 achieves state-of-the-art performance across a diverse set of tasks and capabilities. When discussing the PaLM 2 family, it is important to distinguish between pre-trained models (of various sizes), fine-tuned variants of these models, and the user-facing products that use these models. In particular, user-facing products typically include additional pre- and post-processing steps. Additionally, the underlying models may evolve over time. Therefore, one should not expect the performance of user-facing products to exactly match the results reported in this report.
翻译:我们介绍了 PaLM 2,一种新的最先进语言模型,相比其前身 PaLM,该模型具有更强的多语言和推理能力,且计算效率更高。PaLM 2 是基于 Transformer 的模型,采用混合目标进行训练。通过对英语和多语言以及推理任务的广泛评估,我们证明 PaLM 2 在不同模型规模的下游任务质量上均有显著提升,同时相较于 PaLM,推理速度更快、效率更高。这种效率提升使得模型能够更广泛地部署,同时响应更迅速,从而实现更自然的交互节奏。PaLM 2 展现出强大的推理能力,其在 BIG-Bench 及其他推理任务上相比 PaLM 有大幅改进。PaLM 2 在一系列负责任的人工智能评估中表现稳定,并支持在推理时控制毒性,无需额外开销或影响其他能力。总体而言,PaLM 2 在多样化的任务和能力中取得了最先进的性能。在讨论 PaLM 2 系列时,需区分预训练模型(不同大小)、这些模型的微调变体,以及使用这些模型的面向用户产品。特别地,面向用户产品通常包含额外的预处理和后处理步骤。此外,底层模型可能随时间演化。因此,不应期望面向用户产品的性能与本报告报告的结果完全吻合。