Collection of genotype data in case-control genetic association studies may often be incomplete for reasons related to genes themselves. This non-ignorable missingness structure, if not appropriately accounted for, can result in participation bias in association analyses. To deal with this issue, Chen et al. (2016) proposed to collect additional genetic information from family members of individuals whose genotype data were not available, and developed a maximum likelihood method for bias correction. In this study, we develop an estimating equation approach to analyzing data collected from this design that allows adjustment of covariates. It jointly estimates odds ratio parameters for genetic association and missingness, where a logistic regression model is used to relate missingness to genotype and other covariates. Our method allows correlation between genotype and covariates while using genetic information from family members to provide information on the missing genotype data. In the estimating equation for genetic association parameters, we weight the contribution of each genotyped subject to the empirical likelihood score function by the inverse probability that the genotype data are available. We evaluate large and finite sample performance of our method via simulation studies and apply it to a family-based case-control study of breast cancer.
翻译:在病例对照遗传关联研究中,基因型数据的收集常因与基因自身相关的原因而不完整。这种不可忽略的缺失结构若未得到恰当处理,可能导致关联分析中的参与偏倚。为解决此问题,Chen等人(2016)提出收集基因型数据缺失个体的家庭成员额外遗传信息,并开发了一种用于偏倚校正的最大似然方法。在本研究中,我们提出一种估计方程方法,用于分析基于该设计收集的数据,该方法允许协变量的调整。它联合估计遗传关联与缺失性的比值比参数,其中使用逻辑回归模型将缺失性与基因型及其他协变量相关联。我们的方法允许基因型与协变量之间存在相关性,同时利用家庭成员的遗传信息为缺失基因型数据提供信息。在遗传关联参数的估计方程中,我们通过基因型数据可获得概率的倒数,对每位已基因分型个体对经验似然得分函数的贡献进行加权。我们通过模拟研究评估了该方法在大样本和有限样本下的性能,并将其应用于一项基于家庭的乳腺癌病例对照研究。