In causal inference, properly selecting the propensity score (PS) model is a popular topic and has been widely investigated in observational studies. In addition, there is a large literature concerning the missing data problem. However, there are very few studies investigating the model selection issue for causal inference when the exposure is missing at random (MAR). In this paper, we discuss how to select both imputation and PS models, which can result in the smallest RMSE of the estimated causal effect. Then, we provide a new criterion, called the ``rank score" for evaluating the overall performance of both models. The simulation studies show that the full imputation plus the outcome-related PS models lead to the smallest RMSE and the rank score can also pick the best models. An application study is conducted to study the causal effect of CVD on the mortality of COVID-19 patients.
翻译:在因果推断中,恰当选择倾向得分(PS)模型是一个热门话题,已在观察性研究中得到广泛探讨。此外,存在大量关于缺失数据问题的文献。然而,当暴露数据随机缺失(MAR)时,针对因果推断的模型选择问题研究甚少。本文讨论了如何选择插补模型与PS模型,以使估计因果效应的均方根误差(RMSE)最小。随后,我们提出了一种称为“秩评分”的新准则,用于评估两种模型的综合性能。模拟研究表明,完全插补结合结局相关的PS模型能产生最小的RMSE,且秩评分亦能筛选出最优模型。一项应用研究被用于探讨心血管疾病(CVD)对COVID-19患者死亡率的因果效应。