A backdoor attack allows a malicious user to manipulate the environment or corrupt the training data, thus inserting a backdoor into the trained agent. Such attacks compromise the RL system's reliability, leading to potentially catastrophic results in various key fields. In contrast, relatively limited research has investigated effective defenses against backdoor attacks in RL. This paper proposes the Recovery Triggered States (RTS) method, a novel approach that effectively protects the victim agents from backdoor attacks. RTS involves building a surrogate network to approximate the dynamics model. Developers can then recover the environment from the triggered state to a clean state, thereby preventing attackers from activating backdoors hidden in the agent by presenting the trigger. When training the surrogate to predict states, we incorporate agent action information to reduce the discrepancy between the actions taken by the agent on predicted states and the actions taken on real states. RTS is the first approach to defend against backdoor attacks in a single-agent setting. Our results show that using RTS, the cumulative reward only decreased by 1.41% under the backdoor attack.


翻译:后门攻击允许恶意用户操纵环境或破坏训练数据,从而将后门植入训练好的智能体中。此类攻击会损害强化学习系统的可靠性,在多个关键领域可能导致灾难性后果。相比之下,针对强化学习后门攻击的有效防御措施研究相对有限。本文提出了一种名为“恢复触发状态(RTS)”的新方法,能够有效保护受害智能体免受后门攻击。RTS方法构建一个替代网络来近似动力学模型。开发者随后可将环境从触发状态恢复至干净状态,从而阻止攻击者通过呈现触发器激活隐藏在智能体中的后门。在训练替代网络预测状态时,我们融入智能体动作信息,以减少智能体在预测状态上的动作与真实状态上的动作之间的差异。RTS是首个在单智能体场景下防御后门攻击的方法。实验结果表明,使用RTS后,在后门攻击下累积奖励仅下降1.41%。

0
下载
关闭预览

相关内容

【CVPR2023】基于强化学习的黑盒模型反演攻击
专知会员服务
24+阅读 · 2023年4月12日
《校准自主性中的信任》2022最新16页slides
专知会员服务
21+阅读 · 2022年12月7日
CVPR2022 | 医学图像分析中基于频率注入的后门攻击
专知会员服务
4+阅读 · 2022年7月9日
专知会员服务
20+阅读 · 2021年8月30日
专知会员服务
24+阅读 · 2021年7月10日
【ICLR2021】神经元注意力蒸馏消除DNN中的后门触发器
专知会员服务
15+阅读 · 2021年1月31日
强化学习最新教程,17页pdf
专知会员服务
182+阅读 · 2019年10月11日
强化学习三篇论文 避免遗忘等
CreateAMind
20+阅读 · 2019年5月24日
逆强化学习-学习人先验的动机
CreateAMind
16+阅读 · 2019年1月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
强化学习族谱
CreateAMind
26+阅读 · 2017年8月2日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2010年12月31日
国家自然科学基金
1+阅读 · 2009年12月31日
Arxiv
0+阅读 · 2023年5月24日
VIP会员
最新内容
分层反无人机系统发展新趋势
专知会员服务
5+阅读 · 9月3日
何为协作武器?
专知会员服务
9+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
13+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
8+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
5+阅读 · 8月31日
《美陆军野战手册(2026年):特种部队》
专知会员服务
8+阅读 · 8月31日
相关VIP内容
【CVPR2023】基于强化学习的黑盒模型反演攻击
专知会员服务
24+阅读 · 2023年4月12日
《校准自主性中的信任》2022最新16页slides
专知会员服务
21+阅读 · 2022年12月7日
CVPR2022 | 医学图像分析中基于频率注入的后门攻击
专知会员服务
4+阅读 · 2022年7月9日
专知会员服务
20+阅读 · 2021年8月30日
专知会员服务
24+阅读 · 2021年7月10日
【ICLR2021】神经元注意力蒸馏消除DNN中的后门触发器
专知会员服务
15+阅读 · 2021年1月31日
强化学习最新教程,17页pdf
专知会员服务
182+阅读 · 2019年10月11日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2010年12月31日
国家自然科学基金
1+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员