Unlearning aims to remove the influence of specific training data sources, but this has proved challenging because the contributions of different sources are entangled within the model. Isolating source contributions to disjoint parameters makes removal easier, though it obstructs joint learning across sources. We propose NULLs (Natively Unlearnable LLMs), a model class that satisfies the two opposing goals of isolating source-specific contributions and learning jointly across sources, by training a set of shared backbone neurons alongside a pool of sparsely activated sinks. During training, information specific to a source naturally concentrates in its sinks while information shared across sources accumulates in the backbone. A source is then unlearned at deployment by disabling its corresponding sinks, with no gradient updates and no access to the retained data. We show that NULLs scales to Wikipedia's ~6M articles, isolating each as an independent source. Unlearning a single article removes knowledge specific to it while preserving facts shared with semantically related articles, closely matching retraining from scratch. We note that unlearning with NULLs is also robust: in a case study of unlearning the Harry Potter books, NULLs resists both adversarial extraction and relearning that reverses post-hoc unlearning. Finally, NULLs preserves general language capabilities, matching a standard transformer on downstream benchmarks. Together, these results suggest that source-level unlearning need not be an afterthought. It can be built natively into LLM training while retaining the benefits of shared representation learning.


翻译:遗忘旨在消除特定训练数据源的影响,但由于不同数据源的贡献在模型中相互纠缠,这一目标实现起来颇具挑战。将各数据源的贡献隔离到不重叠的参数上虽便于移除,但却阻碍了跨数据源的联合学习。我们提出NULLs(原生不可遗忘的大语言模型),这是一类满足以下两个对立目标的模型:隔离特定数据源的贡献,并实现跨数据源的联合学习。其方法是在训练一组共享骨干神经元的同时,配备一个由稀疏激活的“汇池”。在训练过程中,特定于某一数据源的信息自然集中于其对应的汇中,而跨数据源共享的信息则积累于骨干网络中。在部署时,通过禁用某个数据源对应的汇即可实现对该源的遗忘,无需梯度更新,也无需访问保留数据。我们证明,NULLs可扩展到维基百科约600万篇文章,将每篇文章隔离为独立的数据源。遗忘单篇文章能移除其特有的知识,同时保留与语义相关文章共享的事实,其效果与从头重新训练高度一致。我们注意到,基于NULLs的遗忘也具有鲁棒性:在以遗忘《哈利·波特》系列书籍为案例的研究中,NULLs能够抵御对抗性提取以及逆转事后遗忘的重新学习。最后,NULLs保留了通用的语言能力,在下游基准测试中与标准Transformer模型表现相当。综合这些结果,我们得出结论:数据源层面的遗忘不必作为事后补救措施,它可以在保留共享表征学习优势的同时,原生地构建到大语言模型的训练过程中。

0
下载
关闭预览

相关内容

大语言模型持续学习:方法、挑战与机遇
专知会员服务
22+阅读 · 3月16日
不可错过!《大语言模型》课程
专知会员服务
31+阅读 · 2025年4月15日
《大语言模型的数据合成与增强综述》
专知会员服务
44+阅读 · 2024年10月19日
大语言模型的知识冲突:成因、根源与展望
专知会员服务
21+阅读 · 2024年9月23日
大型语言模型对齐技术综述:RLHF、RLAIF、PPO、DPO 等
专知会员服务
55+阅读 · 2024年7月24日
大型语言模型:原理、实现与发展
专知会员服务
102+阅读 · 2023年11月28日
Nat. Med. | 医学中的大型语言模型
专知会员服务
58+阅读 · 2023年9月19日
「知识增强预训练语言模型」最新研究综述
专知
18+阅读 · 2022年11月18日
绝对干货!NLP预训练模型:从transformer到albert
新智元
14+阅读 · 2019年11月10日
一大批中文(BERT等)预训练模型等你认领!
PaperWeekly
15+阅读 · 2019年6月25日
NLP通用模型诞生?一个模型搞定十大自然语言常见任务
人工智能头条
10+阅读 · 2018年6月29日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
18+阅读 · 2023年9月2日
Arxiv
21+阅读 · 2023年7月12日
A Survey of Large Language Models
Arxiv
501+阅读 · 2023年3月31日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
1+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关VIP内容
大语言模型持续学习:方法、挑战与机遇
专知会员服务
22+阅读 · 3月16日
不可错过!《大语言模型》课程
专知会员服务
31+阅读 · 2025年4月15日
《大语言模型的数据合成与增强综述》
专知会员服务
44+阅读 · 2024年10月19日
大语言模型的知识冲突:成因、根源与展望
专知会员服务
21+阅读 · 2024年9月23日
大型语言模型对齐技术综述:RLHF、RLAIF、PPO、DPO 等
专知会员服务
55+阅读 · 2024年7月24日
大型语言模型:原理、实现与发展
专知会员服务
102+阅读 · 2023年11月28日
Nat. Med. | 医学中的大型语言模型
专知会员服务
58+阅读 · 2023年9月19日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员