Online advertising has recently grown into a highly competitive and complex multi-billion-dollar industry, with advertisers bidding for ad slots at large scales and high frequencies. This has resulted in a growing need for efficient "auto-bidding" algorithms that determine the bids for incoming queries to maximize advertisers' targets subject to their specified constraints. This work explores efficient online algorithms for a single value-maximizing advertiser under an increasingly popular constraint: Return-on-Spend (RoS). We quantify efficiency in terms of regret relative to the optimal algorithm, which knows all queries a priori. We contribute a simple online algorithm that achieves near-optimal regret in expectation while always respecting the specified RoS constraint when the input sequence of queries are i.i.d. samples from some distribution. We also integrate our results with the previous work of Balseiro, Lu, and Mirrokni [BLM20] to achieve near-optimal regret while respecting both RoS and fixed budget constraints. Our algorithm follows the primal-dual framework and uses online mirror descent (OMD) for the dual updates. However, we need to use a non-canonical setup of OMD, and therefore the classic low-regret guarantee of OMD, which is for the adversarial setting in online learning, no longer holds. Nonetheless, in our case and more generally where low-regret dynamics are applied in algorithm design, the gradients encountered by OMD can be far from adversarial but influenced by our algorithmic choices. We exploit this key insight to show our OMD setup achieves low regret in the realm of our algorithm.
翻译:在线广告已迅速发展为一个高度竞争且复杂的数十亿美元产业,广告主以大规模和高频率竞价广告位。这催生了对高效"自动竞价"算法的迫切需求,该算法需在满足特定约束的前提下,为每次查询确定竞价以最大化广告主目标。本文研究单一价值最大化广告主在日益流行的"投资回报率"约束下的高效在线算法。我们以相对于能预知所有查询的最优算法的遗憾值来量化效率。当输入查询序列独立同分布于某分布时,我们提出一种简单在线算法,能在始终满足指定投资回报率约束的情况下,实现与最优期望遗憾几乎无差的效果。我们还将结果与Balseiro、Lu和Mirrokni [BLM20]的先前工作整合,在同时满足投资回报率与固定预算约束时达到近最优遗憾。该算法采用原始-对偶框架,并使用在线镜像下降法(OMD)进行对偶更新。然而,我们需采用非规范的OMD设置,因此OMD在在线学习对抗性场景下的经典低遗憾保证不再成立。但在此类更普遍的将低遗憾动态应用于算法设计的情境中,OMD所处理的梯度远非对抗性,而是受我们算法选择的影响。我们利用这一关键洞察,证明在该算法框架下我们的OMD设置能实现低遗憾。