Online experimentation with interference is a common challenge in modern applications such as e-commerce and adaptive clinical trials in medicine. For example, in online marketplaces, the revenue of a good depends on discounts applied to competing goods. Statistical inference with interference is widely studied in the offline setting, but far less is known about how to adaptively assign treatments to minimize regret. We address this gap by studying a multi-armed bandit (MAB) problem where a learner (e-commerce platform) sequentially assigns one of possible $\mathcal{A}$ actions (discounts) to $N$ units (goods) over $T$ rounds to minimize regret (maximize revenue). Unlike traditional MAB problems, the reward of each unit depends on the treatments assigned to other units, i.e., there is interference across the underlying network of units. With $\mathcal{A}$ actions and $N$ units, minimizing regret is combinatorially difficult since the action space grows as $\mathcal{A}^N$. To overcome this issue, we study a sparse network interference model, where the reward of a unit is only affected by the treatments assigned to $s$ neighboring units. We use tools from discrete Fourier analysis to develop a sparse linear representation of the unit-specific reward $r_n: [\mathcal{A}]^N \rightarrow \mathbb{R} $, and propose simple, linear regression-based algorithms to minimize regret. Importantly, our algorithms achieve provably low regret both when the learner observes the interference neighborhood for all units and when it is unknown. This significantly generalizes other works on this topic which impose strict conditions on the strength of interference on a known network, and also compare regret to a markedly weaker optimal action. Empirically, we corroborate our theoretical findings via numerical simulations.
翻译:在线实验中的干扰效应是现代应用中普遍存在的挑战,例如电子商务和医学领域的适应性临床试验。以在线市场为例,商品的收益往往取决于对其竞争商品所施加的折扣。尽管在离线环境下,针对干扰的统计推断已得到广泛研究,但关于如何自适应地分配干预措施以最小化遗憾,目前所知甚少。本文通过研究一个多臂老虎机问题来填补这一空白:学习者(电子商务平台)在T轮实验中,依次为N个单元(商品)分配$\mathcal{A}$种可能动作(折扣)之一,以最小化遗憾(即最大化收益)。与传统多臂老虎机问题不同,每个单元的收益不仅取决于自身所接受的处理,还受到其他单元处理方式的影响,即单元底层网络间存在干扰。在$\mathcal{A}$个动作与N个单元的情况下,由于动作空间以$\mathcal{A}^N$规模增长,最小化遗憾成为组合复杂度极高的难题。为克服此问题,我们研究了一种稀疏网络干扰模型,其中单元的收益仅受s个相邻单元处理方式的影响。我们利用离散傅里叶分析工具,构建了单元特定收益$r_n: [\mathcal{A}]^N \rightarrow \mathbb{R} $的稀疏线性表示,并提出了基于线性回归的简洁算法以最小化遗憾。值得注意的是,无论学习者是否观测到所有单元的干扰邻域,我们的算法均能实现理论可证的低遗憾度。这显著拓展了该领域的现有研究成果——以往研究往往需对已知网络上的干扰强度施加严格限制,且其遗憾度是相对于明显更弱的最优动作进行衡量的。最后,我们通过数值模拟验证了理论结论。