Editing pretrained neural networks requires specialized algorithms tailored to specific objectives. Designing such algorithms is often time-consuming and demands significant effort. We present an exploratory framework that formulates neural model editing as a reinforcement learning problem, where agents modify models using reward feedback. We introduce two environments: MaskWorld, where agents scale weights multiplicatively, and ShiftWorld, where agents apply additive weight updates. The reward function combines a utility-preservation objective with a task-specific editing objective, enabling agents to learn targeted modifications while maintaining overall model performance. We evaluate the framework on bias mitigation in text classification and machine unlearning in image classification, both of which traditionally rely on specialized algorithms. Our results show that the learned policies reduce forget set accuracy to nearly 0% while preserving over 90% retain set accuracy on the unlearning task. In the bias mitigation setting, the learned policies improve bias-related performance by more than 5% while maintaining general classification utility. Our findings show that neural model editing can be cast as a reinforcement learning problem, allowing editing policies to be learned from reward feedback rather than manually engineered for each task.


翻译:编辑预训练神经网络需要针对特定目标定制的专用算法,这类算法的设计过程通常耗时且工作量巨大。我们提出一个探索性框架,将神经模型编辑形式化为强化学习问题,智能体通过奖励反馈修改模型。我们引入两个环境:MaskWorld(智能体以乘法方式缩放权重)与ShiftWorld(智能体进行加法权重更新)。奖励函数结合了效用保持目标与任务特定编辑目标,使智能体在学习针对性修改的同时维持模型整体性能。我们在文本分类的偏差缓解与图像分类的机器遗忘任务上评估该框架,这两类任务传统上均依赖专用算法。结果表明,在遗忘任务中,学习到的策略可将遗忘集准确率降至接近0%,同时保留集准确率保持在90%以上。在偏差缓解场景中,学习到的策略将偏差相关性能提升超过5%,同时维持通用分类效用。研究证实神经模型编辑可构建为强化学习问题,使编辑策略能通过奖励反馈习得,无需为每个任务手动设计算法。

0
下载
关闭预览

相关内容

【强化学习】深度强化学习初学者指南
专知会员服务
185+阅读 · 2019年12月14日
2019年新书推荐-《神经网络与深度学习》-Michael Nielsen
深度学习与NLP
14+阅读 · 2019年2月21日
展望:模型驱动的深度学习
人工智能学家
12+阅读 · 2018年1月23日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
7+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Arxiv
0+阅读 · 5月4日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
7+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
14+阅读 · 8月7日
相关VIP内容
【强化学习】深度强化学习初学者指南
专知会员服务
185+阅读 · 2019年12月14日
相关基金
国家自然科学基金
7+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员