Discount regularization, using a shorter planning horizon when calculating the optimal policy, is a popular choice to restrict planning to a less complex set of policies when estimating an MDP from sparse or noisy data (Jiang et al., 2015). It is commonly understood that discount regularization functions by de-emphasizing or ignoring delayed effects. In this paper, we reveal an alternate view of discount regularization that exposes unintended consequences. We demonstrate that planning under a lower discount factor produces an identical optimal policy to planning using any prior on the transition matrix that has the same distribution for all states and actions. In fact, it functions like a prior with stronger regularization on state-action pairs with more transition data. This leads to poor performance when the transition matrix is estimated from data sets with uneven amounts of data across state-action pairs. Our equivalence theorem leads to an explicit formula to set regularization parameters locally for individual state-action pairs rather than globally. We demonstrate the failures of discount regularization and how we remedy them using our state-action-specific method across simple empirical examples as well as a medical cancer simulator.
翻译:折扣正则化是一种流行的方法,即在从稀疏或噪声数据估计MDP时,采用更短的规划视界来计算最优策略,从而将规划限制在较简单的策略集合内(Jiang等人,2015)。通常认为,折扣正则化通过弱化或忽略延迟效应来发挥作用。在本文中,我们揭示了折扣正则化的另一种视角,暴露了其意外后果。我们证明,在较低折扣因子下进行规划所产生的最优策略,与在转移矩阵的所有状态和动作上使用相同分布的任意先验进行规划所得到的最优策略完全相同。实际上,它类似于一个对转移数据更多的状态-动作对施加更强正则化的先验。当转移矩阵从各状态-动作对数据量不均匀的数据集中估计时,这会导致性能不佳。我们的等价定理给出了一个显式公式,用于为单个状态-动作对局部设置正则化参数,而非全局统一设置。我们通过简单的实证示例以及一个医疗癌症模拟器,展示了折扣正则化的失败之处,以及如何使用我们提出的状态-动作特定方法进行补救。