The central challenges in missing data models concern the identifiability of two distributions: the target law and the full law. The target law refers to the joint distribution of the data variables, whereas the full law refers to the joint distribution of the data variables and their corresponding response indicators. However, the relationship between the identifiability of these two distributions and the feasibility of multiple imputation has not been clearly established when the data are missing not at random (MNAR). We present a procedure in which the choice of imputation method is guided by identifiability considerations. The identifiability of the full law implies the applicability of conditionally complete imputation methods that draw imputations for all missing data patterns. By contrast, the non-identifiability of the full law implies that any multiple imputation method aiming to implement conditionally complete imputation will produce biased estimates, thereby also restricting the options for estimating the target law. We demonstrate that alternative imputation strategies can sometimes enable the estimation of the target law in such cases. Specifically, we introduce factorizable imputation where certain observed values are also imputed and the imputed data are weighted in the analysis.
翻译:暂无翻译