Most existing works focus on direct perturbations to the victim's state/action or the underlying transition dynamics to demonstrate the vulnerability of reinforcement learning agents to adversarial attacks. However, such direct manipulations may not be always realizable. In this paper, we consider a multi-agent setting where a well-trained victim agent $\nu$ is exploited by an attacker controlling another agent $\alpha$ with an \textit{adversarial policy}. Previous models do not account for the possibility that the attacker may only have partial control over $\alpha$ or that the attack may produce easily detectable "abnormal" behaviors. Furthermore, there is a lack of provably efficient defenses against these adversarial policies. To address these limitations, we introduce a generalized attack framework that has the flexibility to model to what extent the adversary is able to control the agent, and allows the attacker to regulate the state distribution shift and produce stealthier adversarial policies. Moreover, we offer a provably efficient defense with polynomial convergence to the most robust victim policy through adversarial training with timescale separation. This stands in sharp contrast to supervised learning, where adversarial training typically provides only \textit{empirical} defenses. Using the Robosumo competition experiments, we show that our generalized attack formulation results in much stealthier adversarial policies when maintaining the same winning rate as baselines. Additionally, our adversarial training approach yields stable learning dynamics and less exploitable victim policies.
翻译:现有工作大多聚焦于对受害者状态/动作或底层转移动态的直接扰动,以证明强化学习代理易受对抗攻击。然而,此类直接操控并非总是可行的。本文考虑一个多智能体场景:一个训练有素的受害者代理$\nu$被控制另一代理$\alpha$的攻击者利用\textit{对抗策略}所利用。现有模型未能考虑攻击者可能仅能部分控制$\alpha$,或攻击可能产生易于检测的"异常"行为。此外,针对这些对抗策略缺乏可证明高效的防御。为解决这些局限,我们引入一个通用攻击框架,该框架可灵活建模攻击者控制代理的程度,并允许攻击者调节状态分布偏移从而生成更隐蔽的对抗策略。此外,我们通过时间尺度分离的对抗训练,提出一种具有多项式收敛至最强鲁棒受害者策略的可证明高效防御。这与监督学习形成鲜明对比——后者中的对抗训练通常仅提供\textit{经验性}防御。利用Robosumo竞赛实验,我们证明在保持与基线相同胜率时,所提通用攻击公式能生成更隐蔽的对抗策略。同时,我们的对抗训练方法能产生稳定的学习动态及更不易被利用的受害者策略。