This paper studies statistical decisions for dynamic treatment assignment problems. Many policies involve dynamics in their treatment assignments where treatments are sequentially assigned to individuals across multiple stages and the effect of treatment at each stage is usually heterogeneous with respect to the prior treatments, past outcomes, and observed covariates. We consider estimating an optimal dynamic treatment rule that guides the optimal treatment assignment for each individual at each stage based on the individual's history. This paper proposes an empirical welfare maximization approach in a dynamic framework. The approach estimates the optimal dynamic treatment rule using data from an experimental or quasi-experimental study. The paper proposes two estimation methods: one solves the treatment assignment problem at each stage through backward induction, and the other solves the whole dynamic treatment assignment problem simultaneously across all stages. We derive finite-sample upper bounds on worst-case average welfare regrets for the proposed methods and show $1/\sqrt{n}$-minimax convergence rates. We also modify the simultaneous estimation method to incorporate intertemporal budget/capacity constraints.
翻译:本文研究动态处理分配问题中的统计决策。许多政策涉及处理分配的动态过程,其中处理在多个阶段依次分配给个体,且每阶段处理效果通常因先前处理、过往结果及观测协变量而存在异质性。我们考虑估计一项最优动态处理规则,该规则基于个体历史数据为每阶段个体推荐最优处理分配。本文提出一个动态框架下的经验福利最大化方法,该方法利用实验或准实验研究数据估计最优动态处理规则。论文提出两种估计方法:一种通过逆向归纳逐阶段求解处理分配问题,另一种则跨所有阶段同时求解整个动态处理分配问题。我们推导出所提方法在最坏情况下平均福利遗憾的有限样本上界,并证明其具有$1/\sqrt{n}$-极小最大收敛速率。此外,我们还对同时估计方法进行改进,以纳入跨期预算/容量约束。