Randomization, as a key technique in clinical trials, can eliminate sources of bias and produce comparable treatment groups. In randomized experiments, the treatment effect is a parameter of general interest. Researchers have explored the validity of using linear models to estimate the treatment effect and perform covariate adjustment and thus improve the estimation efficiency. However, the relationship between covariates and outcomes is not necessarily linear, and is often intricate. Advances in statistical theory and related computer technology allow us to use nonparametric and machine learning methods to better estimate the relationship between covariates and outcomes and thus obtain further efficiency gains. However, theoretical studies on how to draw valid inferences when using nonparametric and machine learning methods under stratified randomization are yet to be conducted. In this paper, we discuss a unified framework for covariate adjustment and corresponding statistical inference under stratified randomization and present a detailed proof of the validity of using local linear kernel-weighted least squares regression for covariate adjustment in treatment effect estimators as a special case. In the case of high-dimensional data, we additionally propose an algorithm for statistical inference using machine learning methods under stratified randomization, which makes use of sample splitting to alleviate the requirements on the asymptotic properties of machine learning methods. Finally, we compare the performances of treatment effect estimators using different machine learning methods by considering various data generation scenarios, to guide practical research.
翻译:随机化作为临床试验中的关键技术,可消除偏倚来源并产生可比较的治疗组。在随机化实验中,处理效应是普遍关注的参数。研究者已探索使用线性模型估计处理效应并进行协变量调整的有效性,从而提升估计效率。然而,协变量与结局之间的关系未必是线性的,且通常错综复杂。统计理论及相关计算机技术的进步使我们能够利用非参数和机器学习方法更好地估计协变量与结局之间的关系,从而进一步提高效率。然而,在分层随机化下使用非参数和机器学习方法进行有效推断的理论研究尚待开展。本文讨论了分层随机化下协变量调整的统一框架及相应的统计推断方法,并详细证明了在特例中使用局部线性核加权最小二乘回归进行协变量调整于处理效应估计量中的有效性。针对高维数据情形,本文额外提出了一种在分层随机化下利用机器学习方法进行统计推断的算法,该算法通过样本拆分来放宽对机器学习方法渐近性质的要求。最后,我们通过考虑多种数据生成场景,比较了不同机器学习方法下处理效应估计量的性能表现,以指导实践研究。