What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?

Despite the remarkable capabilities of modern large language models (LLMs), the mechanisms behind their problem-solving abilities remain elusive. In this work, we aim to better understand how the learning dynamics of LLM finetuning shapes downstream generalization. Our analysis focuses on reasoning tasks, whose problem structure allows us to distinguish between memorization (the exact replication of reasoning steps from the training data) and performance (the correctness of the final solution). We find that a model's generalization behavior can be effectively characterized by a training metric we call pre-memorization train accuracy: the accuracy of model samples on training queries before they begin to copy the exact reasoning steps from the training set. On the dataset level, this metric is able to reliably predict test accuracy, achieving $R^2$ of around or exceeding 0.9 across various models (Llama3 8, Gemma2 9B), datasets (GSM8k, MATH), and training configurations. On a per-example level, this metric is also indicative of whether individual model predictions are robust to perturbations in the training query. By connecting a model's learning behavior to its generalization, pre-memorization train accuracy can guide targeted improvements to training strategies. We focus on data curation as an example, and show that prioritizing examples with low pre-memorization accuracy leads to 1.5-2x improvements in data efficiency compared to i.i.d. data scaling, and outperforms other standard data curation techniques.

翻译：尽管现代大型语言模型（LLM）展现出卓越的能力，但其问题解决能力背后的机制仍不明确。本研究旨在深入理解LLM微调的学习动态如何影响下游泛化性能。我们的分析聚焦于推理任务，其问题结构使我们能够区分记忆（精确复现训练数据中的推理步骤）与性能（最终解决方案的正确性）。我们发现，模型的泛化行为可以通过一个称为预记忆训练准确率的训练指标有效表征：该指标衡量模型在开始复制训练集中精确推理步骤之前，对训练查询的采样准确率。在数据集层面，该指标能够可靠地预测测试准确率，在不同模型（Llama3 8B、Gemma2 9B）、数据集（GSM8k、MATH）和训练配置下，其$R^2$值达到约0.9或更高。在单样本层面，该指标也能指示个体模型预测是否对训练查询的扰动具有鲁棒性。通过将模型的学习行为与其泛化能力相关联，预记忆训练准确率可以指导训练策略的针对性改进。我们以数据筛选为例，证明优先选择预记忆准确率较低的样本，相较于独立同分布的数据扩展，可将数据效率提升1.5-2倍，且优于其他标准数据筛选技术。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

【ACL2020】多模态信息抽取，365页ppt

专知会员服务

151+阅读 · 2020年7月6日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

35+阅读 · 2019年10月18日