We propose FLARE, the first fingerprinting mechanism to verify whether a suspected Deep Reinforcement Learning (DRL) policy is an illegitimate copy of another (victim) policy. We first show that it is possible to find non-transferable, universal adversarial masks, i.e., perturbations, to generate adversarial examples that can successfully transfer from a victim policy to its modified versions but not to independently trained policies. FLARE employs these masks as fingerprints to verify the true ownership of stolen DRL policies by measuring an action agreement value over states perturbed via such masks. Our empirical evaluations show that FLARE is effective (100% action agreement on stolen copies) and does not falsely accuse independent policies (no false positives). FLARE is also robust to model modification attacks and cannot be easily evaded by more informed adversaries without negatively impacting agent performance. We also show that not all universal adversarial masks are suitable candidates for fingerprints due to the inherent characteristics of DRL policies. The spatio-temporal dynamics of DRL problems and sequential decision-making process make characterizing the decision boundary of DRL policies more difficult, as well as searching for universal masks that capture the geometry of it.
翻译:我们提出 FLARE,这是首个用于验证疑似深度强化学习(DRL)策略是否为其他(受害)策略非法复制品的指纹识别机制。首先,我们证明了可以找到非可迁移的通用对抗掩码(即扰动),用于生成能够成功从受害策略迁移至其修改版本、但无法迁移至独立训练策略的对抗样本。FLARE 利用这些掩码作为指纹,通过测量经此类掩码扰动的状态上的动作一致值,来验证被盗 DRL 策略的真实所有权。我们的实证评估表明,FLARE 对被盗副本具有有效性(动作一致率为100%),且不会误判独立策略(无假阳性)。同时,FLARE 对模型修改攻击具有鲁棒性,且难以被更精明的对手轻易规避(在不影响代理性能的前提下)。我们还发现,由于 DRL 策略的内在特性,并非所有通用对抗掩码都适合用作指纹。DRL 问题的时空动态特性及序列决策过程不仅增加了刻画 DRL 策略决策边界的难度,也加大了搜索能捕捉其几何结构的通用掩码的难度。