Causal weighted quantile treatment effects (WQTE) are a useful compliment to standard causal contrasts that focus on the mean when interest lies at the tails of the counterfactual distribution. To-date, however, methods for estimation and inference regarding causal WQTEs have assumed complete data on all relevant factors. Missing or incomplete data, however, is a widespread challenge in practical settings, particularly when the data are not collected for research purposes such as electronic health records and disease registries. Furthermore, in such settings may be particularly susceptible to the outcome data being missing-not-at-random (MNAR). In this paper, we consider the use of double-sampling, through which the otherwise missing data is ascertained on a sub-sample of study units, as a strategy to mitigate bias due to MNAR data in the estimation of causal WQTEs. With the additional data in-hand, we present identifying conditions that do not require assumptions regarding missingness in the original data. We then propose a novel inverse-probability weighted estimator and derive its' asymptotic properties, both pointwise at specific quantiles and uniform across a range of quantiles in (0,1), when the propensity score and double-sampling probabilities are estimated. For practical inference, we develop a bootstrap method that can be used for both pointwise and uniform inference. A simulation study is conducted to examine the finite sample performance of the proposed estimators.
翻译:因果加权分位数处理效应(WQTE)是对标准因果对比(聚焦于反事实分布均值)的有益补充,尤其当研究兴趣位于分布尾部时。然而,迄今为止,关于因果WQTE的估计与推断方法均假设所有相关因素的数据完整。但在实际场景中,数据缺失或不完整是普遍挑战,尤其当数据并非为研究目的而收集时(如电子健康记录和疾病登记系统)。此外,这类场景下的结局数据可能特别容易发生非随机缺失(MNAR)。本文考虑使用双重抽样策略(即对研究单元的子样本重新获取原本缺失的数据)以减轻因MNAR数据导致的因果WQTE估计偏差。基于额外获取的数据,我们提出了无需对原始数据缺失机制进行假设的可识别条件。随后,我们构建了一种新的逆概率加权估计量,并推导了其渐近性质——包括特定分位点的逐点性质以及(0,1)范围内分位序列的一致性质——当倾向得分和双重抽样概率被估计时。为进行实用推断,我们开发了一种可用于逐点和一致推断的Bootstrap方法。通过模拟研究检验了所提估计量的有限样本性能。