Policy learning utilizing observational data is pivotal across various domains, with the objective of learning the optimal treatment assignment policy while adhering to specific constraints such as fairness, budget, and simplicity. This study introduces a novel positivity-free (stochastic) policy learning framework designed to address the challenges posed by the impracticality of the positivity assumption in real-world scenarios. This framework leverages incremental propensity score policies to adjust propensity score values instead of assigning fixed values to treatments. We characterize these incremental propensity score policies and establish identification conditions, employing semiparametric efficiency theory to propose efficient estimators capable of achieving rapid convergence rates, even when integrated with advanced machine learning algorithms. This paper provides a thorough exploration of the theoretical guarantees associated with policy learning and validates the proposed framework's finite-sample performance through comprehensive numerical experiments, ensuring the identification of causal effects from observational data is both robust and reliable.
翻译:利用观测数据进行策略学习在众多领域中至关重要,其目标是在遵循公平性、预算和简洁性等特定约束条件下,学习最优的治疗分配策略。本研究提出了一种新颖的无正性(随机)策略学习框架,旨在解决现实场景中正性假设不切实际所带来的挑战。该框架利用增量倾向评分策略来调整倾向评分值,而非为治疗方案分配固定值。我们描述了这些增量倾向评分策略的特征,并建立了识别条件,运用半参数效率理论提出高效估计量,即使在结合先进机器学习算法时也能实现快速收敛速度。本文深入探讨了与策略学习相关的理论保证,并通过全面的数值实验验证了所提框架的有限样本性能,确保从观测数据中识别因果效应既稳健又可靠。