Bilateral trade models the problem of facilitating trades between a seller and a buyer having private valuations for the item being sold. In the online version of the problem, the learner faces a new seller and buyer at each time step, and has to post a price for each of the two parties without any knowledge of their valuations. We consider a scenario where, at each time step, before posting prices the learner observes a context vector containing information about the features of the item for sale. The valuations of both the seller and the buyer follow an unknown linear function of the context. In this setting, the learner could leverage previous transactions in an attempt to estimate private valuations. We characterize the regret regimes of different settings, taking as a baseline the best context-dependent prices in hindsight. First, in the setting in which the learner has two-bit feedback and strong budget balance constraints, we propose an algorithm with $O(\log T)$ regret. Then, we study the same set-up with noisy valuations, providing a tight $\widetilde O(T^{\frac23})$ regret upper bound. Finally, we show that loosening budget balance constraints allows the learner to operate under more restrictive feedback. Specifically, we show how to address the one-bit, global budget balance setting through a reduction from the two-bit, strong budget balance setup. This established a fundamental trade-off between the quality of the feedback and the strictness of the budget constraints.
翻译:双边交易模型旨在促进卖方与买方之间的交易,双方对交易物品均持有私有估值。在该问题的在线版本中,学习者在每个时间步面对新的卖方与买方,且必须在不知晓双方估值的情况下分别为其设定价格。我们考虑如下场景:在每个时间步,学习者在设定价格前会观测到一个上下文向量,其中包含待售物品的特征信息。卖方与买方的估值均遵循一个关于该上下文的未知线性函数。在此设定下,学习者可利用历史交易数据来尝试估计私有估值。我们以事后最优的上下文相关价格为基准,对不同设定下的遗憾界进行了刻画。首先,在学习者具有双比特反馈且面临严格预算平衡约束的设定下,我们提出了一种具有 $O(\log T)$ 遗憾的算法。随后,我们在估值存在噪声的情况下研究同一设定,给出了紧致的 $\widetilde O(T^{\frac23})$ 遗憾上界。最后,我们证明放宽预算平衡约束可使学习者在更严格的反馈条件下运行。具体而言,我们展示了如何通过从双比特、强预算平衡设定进行规约,来处理单比特、全局预算平衡设定。这揭示了反馈质量与预算约束严格性之间的基本权衡关系。