Personalized pricing, which involves tailoring prices based on individual characteristics, is commonly used by firms to implement a consumer-specific pricing policy. In this process, buyers can also strategically manipulate their feature data to obtain a lower price, incurring certain manipulation costs. Such strategic behavior can hinder firms from maximizing their profits. In this paper, we study the contextual dynamic pricing problem with strategic buyers. The seller does not observe the buyer's true feature, but a manipulated feature according to buyers' strategic behavior. In addition, the seller does not observe the buyers' valuation of the product, but only a binary response indicating whether a sale happens or not. Recognizing these challenges, we propose a strategic dynamic pricing policy that incorporates the buyers' strategic behavior into the online learning to maximize the seller's cumulative revenue. We first prove that existing non-strategic pricing policies that neglect the buyers' strategic behavior result in a linear $\Omega(T)$ regret with $T$ the total time horizon, indicating that these policies are not better than a random pricing policy. We then establish that our proposed policy achieves a sublinear regret upper bound of $O(\sqrt{T})$. Importantly, our policy is not a mere amalgamation of existing dynamic pricing policies and strategic behavior handling algorithms. Our policy can also accommodate the scenario when the marginal cost of manipulation is unknown in advance. To account for it, we simultaneously estimate the valuation parameter and the cost parameter in the online pricing policy, which is shown to also achieve an $O(\sqrt{T})$ regret bound. Extensive experiments support our theoretical developments and demonstrate the superior performance of our policy compared to other pricing policies that are unaware of the strategic behaviors.
翻译:个性化定价基于个体特征定制价格,是企业实施消费者差异化定价策略的常见做法。在此过程中,买家可能通过操纵自身特征数据以获取更低价格,并为此承担相应操纵成本。此类战略行为会阻碍企业实现利润最大化。本文研究了存在战略买家情境下的动态定价问题。卖家无法观测到买家的真实特征,只能观察到经买家战略行为操纵后的特征数据。此外,卖家仅能获取表示交易是否达成的二元响应信息,而无法获知买家对产品的估值。针对上述挑战,我们提出一种将买家战略行为纳入在线学习框架的战略动态定价策略,以最大化卖家的累积收益。首先证明,忽视买家战略行为的现有非战略定价策略会导致线性遗憾界限Ω(T)(其中T为总时间跨度),表明此类策略并不优于随机定价策略。进一步证明,我们提出的策略可实现次线性遗憾上界O(√T)。值得注意的是,该策略并非现有动态定价策略与战略行为处理算法的简单组合。当操纵边际成本事先未知时,我们的策略依然适用。为此,我们在在线定价策略中同步估计估值参数与成本参数,并证明该方法同样能达到O(√T)的遗憾界限。大量实验验证了理论推导结果,并表明与未考虑战略行为的其他定价策略相比,本策略具有显著优越性。