Bandits play a crucial role in interactive learning schemes and modern recommender systems. However, these systems often rely on sensitive user data, making privacy a critical concern. This paper investigates privacy in bandits with a trusted centralized decision-maker through the lens of interactive Differential Privacy (DP). While bandits under pure $\epsilon$-global DP have been well-studied, we contribute to the understanding of bandits under zero Concentrated DP (zCDP). We provide minimax and problem-dependent lower bounds on regret for finite-armed and linear bandits, which quantify the cost of $\rho$-global zCDP in these settings. These lower bounds reveal two hardness regimes based on the privacy budget $\rho$ and suggest that $\rho$-global zCDP incurs less regret than pure $\epsilon$-global DP. We propose two $\rho$-global zCDP bandit algorithms, AdaC-UCB and AdaC-GOPE, for finite-armed and linear bandits respectively. Both algorithms use a common recipe of Gaussian mechanism and adaptive episodes. We analyze the regret of these algorithms to show that AdaC-UCB achieves the problem-dependent regret lower bound up to multiplicative constants, while AdaC-GOPE achieves the minimax regret lower bound up to poly-logarithmic factors. Finally, we provide experimental validation of our theoretical results under different settings.
翻译:Bandit算法在交互式学习机制和现代推荐系统中扮演着关键角色。然而,这些系统通常依赖敏感的用户数据,使得隐私保护成为关键议题。本文通过交互式差分隐私(DP)的视角,研究具有可信集中决策者的Bandit隐私问题。虽然纯ε-全局DP下的Bandit已被充分研究,我们致力于深入理解零集中差分隐私(zCDP)下的Bandit。针对有限臂线性Bandit,我们给出了遗憾的极小极大和问题依赖下界,量化了ρ-全局zCDP在这些场景中的代价。这些下界揭示了基于隐私预算ρ的两个难度区间,并表明ρ-全局zCDP产生的遗憾低于纯ε-全局DP。我们提出了两种ρ-全局zCDP Bandit算法:AdaC-UCB(适用于有限臂)和AdaC-GOPE(适用于线性臂)。两种算法均采用高斯机制与自适应回合的通用框架。通过算法遗憾分析,我们证明AdaC-UCB在乘法常数范围内达到问题依赖的遗憾下界,而AdaC-GOPE在对数多项式因子范围内达到极小极大遗憾下界。最后,我们在不同设置下通过实验验证了理论结果。