We introduce PaLM 2, a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM. PaLM 2 is a Transformer-based model trained using a mixture of objectives. Through extensive evaluations on English and multilingual language, and reasoning tasks, we demonstrate that PaLM 2 has significantly improved quality on downstream tasks across different model sizes, while simultaneously exhibiting faster and more efficient inference compared to PaLM. This improved efficiency enables broader deployment while also allowing the model to respond faster, for a more natural pace of interaction. PaLM 2 demonstrates robust reasoning capabilities exemplified by large improvements over PaLM on BIG-Bench and other reasoning tasks. PaLM 2 exhibits stable performance on a suite of responsible AI evaluations, and enables inference-time control over toxicity without additional overhead or impact on other capabilities. Overall, PaLM 2 achieves state-of-the-art performance across a diverse set of tasks and capabilities. When discussing the PaLM 2 family, it is important to distinguish between pre-trained models (of various sizes), fine-tuned variants of these models, and the user-facing products that use these models. In particular, user-facing products typically include additional pre- and post-processing steps. Additionally, the underlying models may evolve over time. Therefore, one should not expect the performance of user-facing products to exactly match the results reported in this report.
翻译:我们介绍了PaLM 2,一种新的最先进语言模型,其多语言与推理能力优于前代PaLM,且计算效率更高。PaLM 2是基于Transformer的模型,采用混合目标函数进行训练。通过对英语、多语言及推理任务的广泛评估,我们证明PaLM 2在不同模型规模下均显著提升了下游任务质量,同时推理速度更快、效率更高,优于PaLM。效率提升不仅支持更广泛的部署,还能实现更快的模型响应,从而促进更自然的交互节奏。PaLM 2展现出强大的推理能力,尤其在BIG-Bench及其他推理任务上较PaLM取得大幅进步。PaLM 2在一系列负责任AI评估中表现稳定,且无需额外开销或影响其他能力即可实现推理时毒性控制。总体而言,PaLM 2在多样化任务与能力维度均达到最先进水平。讨论PaLM 2系列时,需明确区分预训练模型(多种规模)、微调变体及使用这些模型的用户端产品。特别地,用户端产品通常包含额外的前后处理步骤,且底层模型可能随时间演化。因此,用户端产品的性能不应期望与本报告所报告结果完全一致。