Stochastic linear contextual bandit algorithms have substantial applications in practice, such as recommender systems, online advertising, clinical trials, etc. Recent works show that optimal bandit algorithms are vulnerable to adversarial attacks and can fail completely in the presence of attacks. Existing robust bandit algorithms only work for the non-contextual setting under the attack of rewards and cannot improve the robustness in the general and popular contextual bandit environment. In addition, none of the existing methods can defend against attacked context. In this work, we provide the first robust bandit algorithm for stochastic linear contextual bandit setting under a fully adaptive and omniscient attack with sub-linear regret. Our algorithm not only works under the attack of rewards, but also under attacked context. Moreover, it does not need any information about the attack budget or the particular form of the attack. We provide theoretical guarantees for our proposed algorithm and show by experiments that our proposed algorithm improves the robustness against various kinds of popular attacks.
翻译:随机线性上下文赌博机算法在推荐系统、在线广告、临床试验等实际应用中具有重要价值。近期研究表明,最优赌博机算法易受对抗攻击,在攻击下可能完全失效。现有鲁棒赌博机算法仅适用于奖励攻击下的非上下文场景,无法提升通用且流行的上下文赌博机环境下的鲁棒性。此外,现有方法均无法防御上下文被攻击的情况。本文首次提出一种针对随机线性上下文赌博机设置的鲁棒赌博机算法,该算法能应对完全自适应且全知的次线性遗憾攻击。我们的算法不仅能在奖励攻击下工作,还能在上下文被攻击时正常运行。更重要的是,它无需任何关于攻击预算或攻击具体形式的信息。我们为所提算法提供了理论保障,并通过实验证明,该算法能有效提升对各类常见攻击的鲁棒性。