Sequential posted pricing auctions are popular because of their simplicity in practice and their tractability in theory. A usual assumption in their study is that the Bayesian prior distributions of the buyers are known to the seller, while in reality these priors can only be accessed from historical data. To overcome this assumption, we study sequential posted pricing in the bandit learning model, where the seller interacts with $n$ buyers over $T$ rounds: In each round the seller posts $n$ prices for the $n$ buyers and the first buyer with a valuation higher than the price takes the item. The only feedback that the seller receives in each round is the revenue. Our main results obtain nearly-optimal regret bounds for single-item sequential posted pricing in the bandit learning model. In particular, we achieve an $\tilde{O}(\mathsf{poly}(n)\sqrt{T})$ regret for buyers with (Myerson's) regular distributions and an $\tilde{O}(\mathsf{poly}(n)T^{{2}/{3}})$ regret for buyers with general distributions, both of which are tight in the number of rounds $T$. Our result for regular distributions was previously not known even for the single-buyer setting and relies on a new half-concavity property of the revenue function in the value space. For $n$ sequential buyers, our technique is to run a generalized single-buyer algorithm for all the buyers and to carefully bound the regret from the sub-optimal pricing of the suffix buyers.
翻译:序列定价拍卖因其在实际应用中的简洁性和理论分析上的可处理性而广受欢迎。对其研究通常假设卖方已知买方的贝叶斯先验分布,然而现实中这些先验分布只能通过历史数据获取。为突破这一假设,我们在老虎机学习模型中研究序列定价问题:卖方在 $T$ 轮中与 $n$ 个买方进行交互,每轮卖方为 $n$ 个买方分别设定价格,首个估值高于定价的买方获得物品。卖方每轮仅能获得收益反馈。我们的主要成果是在老虎机学习模型中为单物品序列定价问题取得了近乎最优的遗憾界。具体而言,对于满足(迈尔森)正则分布的买方,我们实现了 $\tilde{O}(\mathsf{poly}(n)\sqrt{T})$ 的遗憾;对于满足一般分布的买方,我们实现了 $\tilde{O}(\mathsf{poly}(n)T^{{2}/{3}})$ 的遗憾,两者在回合数 $T$ 上均达到紧界。其中针对正则分布的结果此前即使在单买方场景下亦属未知,其证明依赖于收益函数在价值空间中具备的新性质——半凹性。对于 $n$ 个序列买方,我们的技术方案是为所有买方运行广义单买方算法,并精细地控制因后续买方次优定价所产生的遗憾。