Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI (RAI), which aims to ensure safe and trustworthy responses, remains underexplored in this context. In FedLLM, client-side training data may contain harmful content, resulting in unsafe LLMs that can generate inappropriate responses. Aggregating such models into a global model and redistributing it to clients risks the widespread deployment of unsafe LLMs. To address this, we incorporate two well-established RAI techniques into FedLLM: safety filtering and constitutional AI. Our experiments show that these methods significantly improve LLM safety, achieving over 20% improvement on AdvBench.


翻译:近期研究日益关注利用联邦学习训练大语言模型(LLMs)的范式(即FedLLM)。然而,旨在确保安全可信响应的负责任人工智能(RAI)在此场景中仍待深入探索。在FedLLM中,客户端训练数据可能含有有害内容,导致LLMs生成不当响应;若将这些模型聚合为全局模型并重新分发至客户端,将面临不安全LLMs大规模部署的风险。针对此问题,我们将两种成熟的RAI技术引入FedLLM:安全过滤与宪法人工智能。实验表明,这些方法显著提升LLMs安全性,在AdvBench上实现超过20%的性能改善。

0
下载
关闭预览

相关内容

大语言模型遇见法律人工智能:综述
专知会员服务
27+阅读 · 2025年9月15日
可解释人工智能中的大语言模型:全面综述
专知会员服务
54+阅读 · 2025年4月2日
大语言模型安全开发者手册:构建安全的 AI 应用程序
专知会员服务
36+阅读 · 2024年9月29日
生成式人工智能大型语言模型的安全性:概述
专知会员服务
36+阅读 · 2024年7月30日
迈向可信的人工智能:伦理和稳健的大型语言模型综述
专知会员服务
40+阅读 · 2024年7月28日
大型语言模型在国家安全应用中的使用
专知会员服务
57+阅读 · 2024年7月13日
大型语言模型网络安全综述
专知会员服务
68+阅读 · 2024年5月12日
联邦学习安全与隐私保护研究综述
专知
12+阅读 · 2020年8月7日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
Arxiv
12+阅读 · 2023年5月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
1+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
4+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
10+阅读 · 8月7日
相关VIP内容
大语言模型遇见法律人工智能:综述
专知会员服务
27+阅读 · 2025年9月15日
可解释人工智能中的大语言模型:全面综述
专知会员服务
54+阅读 · 2025年4月2日
大语言模型安全开发者手册:构建安全的 AI 应用程序
专知会员服务
36+阅读 · 2024年9月29日
生成式人工智能大型语言模型的安全性:概述
专知会员服务
36+阅读 · 2024年7月30日
迈向可信的人工智能:伦理和稳健的大型语言模型综述
专知会员服务
40+阅读 · 2024年7月28日
大型语言模型在国家安全应用中的使用
专知会员服务
57+阅读 · 2024年7月13日
大型语言模型网络安全综述
专知会员服务
68+阅读 · 2024年5月12日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
Top
微信扫码咨询专知VIP会员