In cluster-randomized trials (CRTs), missing data can occur in various ways, including missing values in outcomes and baseline covariates at the individual or cluster level, or completely missing information for non-participants. Among the various types of missing data in CRTs, missing outcomes have attracted the most attention. However, no existing methods can simultaneously address all aforementioned types of missing data in CRTs. To fill in this gap, we propose a new doubly-robust estimator for the average treatment effect on a variety of scales. The proposed estimator simultaneously handles missing outcomes under missingness at random, missing covariates without constraining the missingness mechanism, and missing cluster-population sizes via a uniform sampling mechanism. Furthermore, we detail key considerations to improve precision by specifying the optimal weights, leveraging machine learning, and modeling the treatment assignment mechanism. Finally, to evaluate the impact of violating missing data assumptions, we contribute a new sensitivity analysis framework tailored to CRTs. Simulation studies and a real data application both demonstrate that our proposed methods are effective in handling missing data in CRTs and superior to the existing methods.
翻译:在整群随机试验(CRTs)中,缺失数据可能以多种形式出现,包括个体或群组层面结局与基线协变量的缺失值,以及非参与者的信息完全缺失。在CRT各类缺失数据中,结局缺失所受关注最多,但现有方法均无法同时处理上述所有类型的CRT缺失数据。为填补这一空白,我们提出一种新型的双稳健估计量,可用于估计多种尺度下的平均处理效应。该估计量能同时处理:随机缺失机制下的结局缺失、无缺失机制约束的协变量缺失,以及通过统一抽样机制处理的群组总体规模缺失。此外,我们详细阐述了通过指定最优权重、利用机器学习及建模处理分配机制提升估计精度的关键考量。最后,为评估违反缺失数据假设的影响,我们构建了专用于CRT的新型敏感性分析框架。模拟研究与实际数据应用均表明:所提出方法能有效处理CRT中的缺失数据,且优于现有方法。