Autonomous agents deployed in the real world need to be robust against adversarial attacks on sensory inputs. Robustifying agent policies requires anticipating the strongest attacks possible. We demonstrate that existing observation-space attacks on reinforcement learning agents have a common weakness: while effective, their lack of information-theoretic detectability constraints makes them detectable using automated means or human inspection. Detectability is undesirable to adversaries as it may trigger security escalations. We introduce {\epsilon}-illusory, a novel form of adversarial attack on sequential decision-makers that is both effective and of {\epsilon}-bounded statistical detectability. We propose a novel dual ascent algorithm to learn such attacks end-to-end. Compared to existing attacks, we empirically find {\epsilon}-illusory to be significantly harder to detect with automated methods, and a small study with human participants (IRB approval under reference R84123/RE001) suggests they are similarly harder to detect for humans. Our findings suggest the need for better anomaly detectors, as well as effective hardware- and system-level defenses. The project website can be found at https://tinyurl.com/illusory-attacks.
翻译:部署在现实世界中的自主智能体需要能够抵御针对感官输入的对抗性攻击。强化智能体策略需要预判可能的最强攻击。我们证明,现有的针对强化学习智能体的观测空间攻击存在一个共同弱点:尽管它们有效,但由于缺乏信息论可检测性约束,这些攻击可通过自动化手段或人工检查被检测出来。可检测性对攻击者而言是不利的,因为它可能触发安全升级。我们提出了一种针对序列决策者的新型对抗攻击形式——{\epsilon}-幻象,它既有效又具有{\epsilon}有界统计可检测性。我们提出了一种新颖的对偶上升算法来端到端地学习此类攻击。与现有攻击相比,我们的实验发现{\epsilon}-幻象更难被自动化方法检测,而一项包含人类参与者的小型研究(伦理审查批准号R84123/RE001)表明,人类同样更难察觉这些攻击。我们的研究结果表明,需要更优秀的异常检测器以及有效的硬件和系统级防御措施。项目网站地址为https://tinyurl.com/illusory-attacks。