We study an online joint assortment-inventory optimization problem, in which we assume that the choice behavior of each customer follows the Multinomial Logit (MNL) choice model, and the attraction parameters are unknown a priori. The retailer makes periodic assortment and inventory decisions to dynamically learn from the realized demands about the attraction parameters while maximizing the expected total profit over time. In this paper, we propose a novel algorithm that can effectively balance the exploration and exploitation in the online decision-making of assortment and inventory. Our algorithm builds on a new estimator for the MNL attraction parameters, a novel approach to incentivize exploration by adaptively tuning certain known and unknown parameters, and an optimization oracle to static single-cycle assortment-inventory planning problems with given parameters. We establish a regret upper bound for our algorithm and a lower bound for the online joint assortment-inventory optimization problem, suggesting that our algorithm achieves nearly optimal regret rate, provided that the static optimization oracle is exact. Then we incorporate more practical approximate static optimization oracles into our algorithm, and bound from above the impact of static optimization errors on the regret of our algorithm. At last, we perform numerical studies to demonstrate the effectiveness of our proposed algorithm.
翻译:本文研究了在线联合分类-库存优化问题,其中假设每位顾客的选择行为服从多项逻辑模型(Multinomial Logit, MNL),且产品的吸引力参数先验未知。零售商通过周期性调整分类与库存决策,从已实现的需求中动态学习吸引力参数,同时最大化长期期望总收益。我们提出一种新型算法,能有效平衡在线分类与库存决策中的探索与利用。该算法基于以下三点构建:针对MNL吸引力参数的新估计器、通过自适应调节已知与未知参数以激励探索的创新方法,以及针对给定参数下静态单周期分类-库存规划问题的优化求解器。我们证明了算法的遗憾上界以及在线联合分类-库存优化问题的下界,表明在静态优化求解器精确的前提下,该算法可达到近乎最优的遗憾率。随后,我们将更实用的近似静态优化求解器融入算法,并从上限角度量化静态优化误差对算法遗憾值的影响。最后通过数值实验验证了所提算法的有效性。