Negative control variables are sometimes used in non-experimental studies to detect the presence of confounding by hidden factors. An outcome is said to be a valid negative control outcome (NCO) or more broadly, an outcome that is a proxy for confounding to the extent that it is influenced by unobserved confounders of the exposure effects on the outcome in view, although not causally impacted by the exposure. Tchetgen Tchetgen (2013) introduced the control outcome calibration approach (COCA), as a formal NCO counterfactual method to detect and correct for residual confounding bias. For identification, COCA treats the NCO as an error-prone proxy of the treatment-free counterfactual outcome of interest, and involves regressing the NCO, on the treatment-free counterfactual, together with a rank-preserving structural model which assumes a constant individual-level causal effect. In this work, we establish nonparametric COCA identification for the average causal effect for the treated, without requiring rank-preservation, therefore accommodating unrestricted effect heterogeneity across units. This nonparametric identification result has important practical implications, as it provides single proxy confounding control, in contrast to recently proposed proximal causal inference, which relies for identification on a pair of confounding proxies. For COCA estimation we propose three separate strategies: (i) an extended propensity score approach, (ii) an outcome bridge function approach, and (iii) a doubly robust approach which is unbiased if either (i) or (ii) is unbiased. Finally, we illustrate the proposed methods in an application evaluating the causal impact of a Zika virus outbreak on birth rate in Brazil.
翻译:负对照变量有时被用于非实验研究,以检测隐藏因素引起的混杂效应。若某个结果变量虽不受暴露因素的因果影响,但受暴露效应中未观测混杂因素的影响,则该结果被视为有效的负对照结果(NCO),或更广义地,作为混杂因素的代理变量。Tchetgen Tchetgen(2013)提出了对照结果校准法(COCA),这是一种正式的NCO反事实方法,用于检测和纠正残余混杂偏倚。在识别过程中,COCA将NCO视为感兴趣的无治疗反事实结果的有误差代理,并通过秩保持结构模型(假设个体因果效应恒定)将NCO对无治疗反事实进行回归分析。本研究在无需秩保持假设的条件下,建立了接受治疗群体的平均因果效应的非参数COCA识别方法,从而允许跨单元的无限制效应异质性。该非参数识别结果具有重要实践意义,因为它提供了单代理混杂控制方法,而不同于近期提出的依赖一对混杂代理进行识别的近端因果推断方法。针对COCA估计,我们提出三种独立策略:(i)扩展倾向性评分方法,(ii)结果桥函数方法,以及(iii)双稳健方法——当方法(i)或(ii)无偏时,该方法亦保持无偏。最后,我们通过评估寨卡病毒爆发对巴西出生率的因果影响案例,对所提方法进行了实证说明。