Contextual dynamic pricing aims to set personalized prices based on sequential interactions with customers. At each time period, a customer who is interested in purchasing a product comes to the platform. The customer's valuation for the product is a linear function of contexts, including product and customer features, plus some random market noise. The seller does not observe the customer's true valuation, but instead needs to learn the valuation by leveraging contextual information and historical binary purchase feedbacks. Existing models typically assume full or partial knowledge of the random noise distribution. In this paper, we consider contextual dynamic pricing with unknown random noise in the valuation model. Our distribution-free pricing policy learns both the contextual function and the market noise simultaneously. A key ingredient of our method is a novel perturbed linear bandit framework, where a modified linear upper confidence bound algorithm is proposed to balance the exploration of market noise and the exploitation of the current knowledge for better pricing. We establish the regret upper bound and a matching lower bound of our policy in the perturbed linear bandit framework and prove a sub-linear regret bound in the considered pricing problem. Finally, we demonstrate the superior performance of our policy on simulations and a real-life auto-loan dataset.
翻译:上下文动态定价旨在基于与客户的顺序交互设定个性化价格。在每个时间周期,有购买产品意向的客户进入平台。客户对产品的估价是上下文的线性函数(包括产品和客户特征)加上随机市场噪声。卖方无法观测到客户真实估价,而需利用上下文信息和历史二元购买反馈来学习估价。现有模型通常假设完全或部分已知随机噪声分布。本文研究了估价模型中随机噪声未知情况下的上下文动态定价问题。我们的无分布假设定价策略同时学习上下文函数和市场噪声。该方法的核心创新在于提出了一种新型扰动线性老虎机框架,通过改进的线性置信上界算法来平衡市场噪声探索与当前知识利用,以实现更优定价。我们在扰动线性老虎机框架中建立了策略的遗憾上界及其匹配下界,并在所考虑的定价问题中证明了亚线性遗憾界。最后,通过仿真实验和真实汽车贷款数据集验证了策略的优越性能。