In this paper we address the problem of learning and backtesting inventory control policies in the presence of general arrival dynamics -- which we term as a quantity-over-time arrivals model (QOT). We also allow for order quantities to be modified as a post-processing step to meet vendor constraints such as order minimum and batch size constraints -- a common practice in real supply chains. To the best of our knowledge this is the first work to handle either arbitrary arrival dynamics or an arbitrary downstream post-processing of order quantities. Building upon recent work (Madeka et al., 2022) we similarly formulate the periodic review inventory control problem as an exogenous decision process, where most of the state is outside the control of the agent. Madeka et al., 2022 show how to construct a simulator that replays historic data to solve this class of problem. In our case, we incorporate a deep generative model for the arrivals process as part of the history replay. By formulating the problem as an exogenous decision process, we can apply results from Madeka et al., 2022 to obtain a reduction to supervised learning. Via simulation studies we show that this approach yields statistically significant improvements in profitability over production baselines. Using data from a real-world A/B test, we show that Gen-QOT generalizes well to off-policy data and that the resulting buying policy outperforms traditional inventory management systems in real world settings.
翻译:本文研究了在通用到达动态下学习与回测库存控制策略的问题——我们将其定义为随时间变化的数量到达模型(QOT)。我们还允许将订单数量作为后处理步骤进行调整,以满足供应商约束(如最小订单量和批量约束),这在实际供应链中是一种常见做法。据我们所知,这是首项能够处理任意到达动态或订单数量任意下游后处理的工作。基于近期研究(Madeka等人,2022),我们同样将定期盘点库存控制问题表述为外生决策过程,其中大部分状态不受智能体控制。Madeka等人(2022)展示了如何构建一个利用历史数据进行回放的模拟器来解决此类问题。在我们的工作中,我们将到达过程的深度生成模型作为历史回放的一部分。通过将问题表述为外生决策过程,我们可应用Madeka等人(2022)的结果,将其简化为监督学习问题。仿真研究表明,该方法在生产基线基础上实现了统计显著的盈利提升。基于真实世界A/B测试的数据,我们证明Gen-QOT对离策略数据具有良好的泛化能力,且由此产生的采购策略在真实场景中优于传统库存管理系统。