We study a problem of designing replication-proof bandit mechanisms when agents strategically register or replicate their own arms to maximize their payoff. We consider Bayesian agents who are unaware of ex-post realization of their own arms' mean rewards, which is the first to study Bayesian extension of Shin et al. (2022). This extension presents significant challenges in analyzing equilibrium, in contrast to the fully-informed setting by Shin et al. (2022) under which the problem simply reduces to a case where each agent only has a single arm. With Bayesian agents, even in a single-agent setting, analyzing the replication-proofness of an algorithm becomes complicated. Remarkably, we first show that the algorithm proposed by Shin et al. (2022), defined H-UCB, is no longer replication-proof for any exploration parameters. Then, we provide sufficient and necessary conditions for an algorithm to be replication-proof in the single-agent setting. These results centers around several analytical results in comparing the expected regret of multiple bandit instances, which might be of independent interest. We further prove that exploration-then-commit (ETC) algorithm satisfies these properties, whereas UCB does not, which in fact leads to the failure of being replication-proof. We expand this result to multi-agent setting, and provide a replication-proof algorithm for any problem instance. The proof mainly relies on the single-agent result, as well as some structural properties of ETC and the novel introduction of a restarting round, which largely simplifies the analysis while maintaining the regret unchanged (up to polylogarithmic factor). We finalize our result by proving its sublinear regret upper bound, which matches that of H-UCB.
翻译:我们研究了当策略性代理人通过注册或复制自己的臂来最大化收益时,设计防复制赌臂机制的问题。我们考虑贝叶斯代理人,他们不知道自身臂平均收益的事后实现,这是首次研究Shin等人(2022)工作的贝叶斯扩展。与Shin等人(2022)完全知情设定下问题简化为每个代理人仅有一个臂的情形相反,这种扩展在均衡分析中带来了显著挑战。在贝叶斯代理人情形下,即使仅在单代理人设定中,分析算法的防复制性也变得复杂。值得注意的是,我们首先证明Shin等人(2022)提出的H-UCB算法对任意探索参数均不再具有防复制性。随后,我们给出了单代理人设定下算法满足防复制性的充分必要条件。这些结果围绕多个赌臂实例期望遗憾比较的若干分析结果展开,这些分析结果可能具有独立意义。我们进一步证明了探索-然后-承诺(ETC)算法满足这些性质,而UCB算法则不然——事实上这导致了其无法实现防复制。我们将该结果推广至多代理人设定,并针对任意问题实例给出了防复制算法。该证明主要依赖于单代理人结果,以及ETC的某些结构属性和重新启动轮次的新颖引入,这在保持遗憾不变(至多相差多对数因子)的同时极大简化了分析。我们通过证明其次线性遗憾上界(与H-UCB相当)来完善最终结果。