Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation. However, their immense number of parameters and complex transformer-based architectures result in significant resource demands and computational complexity during training, making it challenging to optimize them efficiently on large datasets. To reduce training costs while preserving performance, researchers have investigated coreset selection techniques, which aim to identify small, representative subsets of the entire training dataset to accelerate LLM training. However, existing coreset selection methods fail to adapt to the dynamic nature of LLM training and often struggle with scalability for models of this size. To address these limitations, we propose a graph-guided adaptive and dynamic coreset selection framework for LLMs, namely GRACE. GRACE dynamically constructs and updates coresets by combining representation diversity with gradient-based importance metrics, ensuring both informativeness and efficiency. To mitigate the computational cost of frequent updates, GRACE leverages a $k$-NN graph-based propagation mechanism and selectively updates scores and embeddings, adapting to evolving training dynamics. Extensive experiments on three benchmarks demonstrate that GRACE significantly improves training efficiency and downstream performance across diverse LLMs and tasks.


翻译:大语言模型在自然语言理解与生成方面展现了卓越能力。然而,其庞大的参数量与基于Transformer的复杂架构导致训练过程中需要大量资源与计算复杂度,使得在大规模数据集上高效优化面临挑战。为在降低训练成本的同时保持性能,研究者探索了核心集选择技术,旨在从完整训练数据中识别小型代表性子集以加速大语言模型训练。然而,现有核心集选择方法难以适应大语言模型训练的动态特性,且在处理此类规模模型时存在扩展性瓶颈。针对上述局限,我们提出面向大语言模型的图引导自适应动态核心集选择框架GRACE。GRACE通过结合表示多样性与基于梯度的重要性指标,动态构建并更新核心集,确保信息性与效率。为缓解频繁更新带来的计算开销,GRACE利用k近邻图传播机制,选择性更新分数与嵌入表示以适应训练动态变化。在三个基准上的广泛实验表明,GRACE能够显著提升不同大语言模型及其任务场景下的训练效率与下游性能。

0
下载
关闭预览

相关内容

赋能大型语言模型多领域资源挑战
专知会员服务
11+阅读 · 2025年6月10日
《大语言模型的数据合成与增强综述》
专知会员服务
44+阅读 · 2024年10月19日
一文速览大语言模型提示最新进展
专知会员服务
82+阅读 · 2023年12月24日
大型语言模型:原理、实现与发展
专知会员服务
102+阅读 · 2023年11月28日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
TextInfoExp:自然语言处理相关实验(基于sougou数据集)
全球人工智能
12+阅读 · 2017年11月12日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员