Sequential posted pricing auctions are popular because of their simplicity in practice and their tractability in theory. A usual assumption in their study is that the Bayesian prior distributions of the buyers are known to the seller, while in reality these priors can only be accessed from historical data. To overcome this assumption, we study sequential posted pricing in the bandit learning model, where the seller interacts with $n$ buyers over $T$ rounds: In each round the seller posts $n$ prices for the $n$ buyers and the first buyer with a valuation higher than the price takes the item. The only feedback that the seller receives in each round is the revenue. Our main results obtain nearly-optimal regret bounds for single-item sequential posted pricing in the bandit learning model. In particular, we achieve an $\tilde{O}(\mathsf{poly}(n)\sqrt{T})$ regret for buyers with (Myerson's) regular distributions and an $\tilde{O}(\mathsf{poly}(n)T^{{2}/{3}})$ regret for buyers with general distributions, both of which are tight in the number of rounds $T$. Our result for regular distributions was previously not known even for the single-buyer setting and relies on a new half-concavity property of the revenue function in the value space. For $n$ sequential buyers, our technique is to run a generalized single-buyer algorithm for all the buyers and to carefully bound the regret from the sub-optimal pricing of the suffix buyers.
翻译:摘要:序贯定价拍卖因其实践中的简便性与理论上的可处理性而广受欢迎。其研究中的常见假设是卖家已知买家的贝叶斯先验分布,然而现实中这些先验只能从历史数据中获取。为克服这一假设,我们研究了强盗学习模型中的序贯定价,其中卖家与$n$个买家在$T$轮中进行交互:每轮中卖家为$n$个买家设定$n$个价格,且首个估价高于价格的买家获得商品。卖家每轮仅能获得的反馈是收益。我们的主要结果在强盗学习模型中针对单品序贯定价获得了近乎最优的遗憾界。具体而言,对于符合(Myerson)正则分布的买家,我们实现了$\tilde{O}(\mathsf{poly}(n)\sqrt{T})$的遗憾值;对于一般分布的买家,实现了$\tilde{O}(\mathsf{poly}(n)T^{{2}/{3}})$的遗憾值,两者在轮数$T$上均为紧界。针对正则分布的结果即使在单买家设置下此前也未被知晓,其关键依赖于收益函数在价值空间中的新半凹性性质。对于$n$个序贯买家,我们的技术是为所有买家运行一个广义的单买家算法,并仔细约束由尾随买家次优定价产生的遗憾。