In studies ranging from clinical medicine to policy research, complete data are usually available from a population $\mathscr{P}$, but the quantity of interest is often sought for a related but different population $\mathscr{Q}$ which only has partial data. In this paper, we consider the setting that both outcome $Y$ and covariate ${\bf X}$ are available from $\mathscr{P}$ whereas only ${\bf X}$ is available from $\mathscr{Q}$, under the so-called label shift assumption, i.e., the conditional distribution of ${\bf X}$ given $Y$ remains the same across the two populations. To estimate the parameter of interest in $\mathscr{Q}$ via leveraging the information from $\mathscr{P}$, the following three ingredients are essential: (a) the common conditional distribution of ${\bf X}$ given $Y$, (b) the regression model of $Y$ given ${\bf X}$ in $\mathscr{P}$, and (c) the density ratio of $Y$ between the two populations. We propose an estimation procedure that only needs standard nonparametric technique to approximate the conditional expectations with respect to (a), while by no means needs an estimate or model for (b) or (c); i.e., doubly flexible to the possible model misspecifications of both (b) and (c). This is conceptually different from the well-known doubly robust estimation in that, double robustness allows at most one model to be misspecified whereas our proposal can allow both (b) and (c) to be misspecified. This is of particular interest in our setting because estimating (c) is difficult, if not impossible, by virtue of the absence of the $Y$-data in $\mathscr{Q}$. Furthermore, even though the estimation of (b) is sometimes off-the-shelf, it can face curse of dimensionality or computational challenges. We develop the large sample theory for the proposed estimator, and examine its finite-sample performance through simulation studies as well as an application to the MIMIC-III database.
翻译:在从临床医学到政策研究等众多领域中,通常可从群体$\mathscr{P}$获得完整数据,但研究者往往需要为仅拥有部分数据的相关但不同群体$\mathscr{Q}$估计目标量。本文考虑如下设定:在标签偏移假设(即给定$Y$后$\bf X$的条件分布在两个群体中保持不变)下,$\mathscr{P}$中同时可获得结果变量$Y$与协变量${\bf X}$,而$\mathscr{Q}$中仅可获得${\bf X}$数据。为通过利用$\mathscr{P}$中的信息估计$\mathscr{Q}$中的目标参数,以下三个要素至关重要:(a) 给定$Y$后$\bf X$的共同条件分布,(b) $\mathscr{P}$中$Y$关于$\bf X$的回归模型,以及(c) 两个群体间$Y$的密度比。我们提出一种估计方法:仅需标准非参数技术近似(a)中的条件期望,而完全无需对(b)或(c)进行估计或建模——即对(b)和(c)的潜在模型误设具有双重灵活性。这与经典的双重稳健估计存在本质区别:双重稳健性最多允许一个模型被误设,而我们的方法可同时允许(b)和(c)被误设。这在当前设定下尤为重要,因为由于$\mathscr{Q}$中缺乏$Y$数据,估计(c)即使可能也十分困难。此外,尽管(b)的估计有时可基于现成方法,但仍可能面临维数灾难或计算挑战。我们建立了所提估计量的大样本理论,并通过模拟研究及对MIMIC-III数据库的应用检验其有限样本表现。