Causal inference is often complicated by post-treatment variables, which appear in many scientific problems, including noncompliance, truncation by death, mediation, and surrogate endpoint evaluation. Principal stratification is a strategy that adjusts for the potential values of the post-treatment variables, defined as the principal strata. It allows for characterizing treatment effect heterogeneity across principal strata and unveiling the mechanism of the treatment on the outcome related to post-treatment variables. However, the existing literature has primarily focused on binary post-treatment variables, leaving the case with continuous post-treatment variables largely unexplored, due to the complexity of infinitely many principal strata that challenge both the identification and estimation of causal effects. We fill this gap by providing nonparametric identification and semiparametric estimation theory for principal stratification with continuous post-treatment variables. We propose to use working models to approximate the underlying causal effect surfaces and derive the efficient influence functions of the corresponding model parameters. Based on the theory, we construct doubly robust estimators and implement them in an R package.
翻译:因果推断常受后处理变量的干扰,这类变量出现在众多科学问题中,包括非依从性、删失截断、中介效应及替代终点评估等。主分层分析通过调整后处理变量的潜在取值(即主分层)来应对这一挑战,既能刻画不同主分层间的处理效应异质性,又可揭示处理通过后处理变量影响结局的机制。然而现有文献主要关注二值后处理变量,由于连续后处理变量存在无穷多个主分层,导致因果效应的识别和估计均面临挑战,相关研究尚属空白。本文填补了这一理论空白,提出连续后处理变量主分层分析的非参数识别理论与半参数估计方法。我们通过工作模型拟合潜在因果效应曲面,推导出相应模型参数的有效影响函数,并基于该理论构建双重稳健估计量,最终将其实现为R语言程序包。