Rates of missing data often depend on record-keeping policies and thus may change across times and locations, even when the underlying features are comparatively stable. In this paper, we introduce the problem of Domain Adaptation under Missingness Shift (DAMS). Here, (labeled) source data and (unlabeled) target data would be exchangeable but for different missing data mechanisms. We show that if missing data indicators are available, DAMS reduces to covariate shift. Addressing cases where such indicators are absent, we establish the following theoretical results for underreporting completely at random: (i) covariate shift is violated (adaptation is required); (ii) the optimal linear source predictor can perform arbitrarily worse on the target domain than always predicting the mean; (iii) the optimal target predictor can be identified, even when the missingness rates themselves are not; and (iv) for linear models, a simple analytic adjustment yields consistent estimates of the optimal target parameters. In experiments on synthetic and semi-synthetic data, we demonstrate the promise of our methods when assumptions hold. Finally, we discuss a rich family of future extensions.
翻译:缺失数据的比率通常取决于记录保存策略,因此即使底层特征相对稳定,缺失率也可能随时间和地点发生变化。本文提出缺失转移下的域自适应问题。在此问题中,(带标签的)源域数据和(未标签的)目标域数据本应可交换,但由于缺失数据机制不同而不可交换。我们证明,若可获得缺失数据指示变量,DAMS可简化为协变量偏移。针对指示变量缺失的情况,我们建立了完全随机漏报情形下的以下理论结果:(i)协变量偏移条件被违反(需进行自适应);(ii)最优线性源预测器在目标域上的性能可能任意差于始终预测均值的方法;(iii)即使在缺失率本身不可识别的情况下,仍可识别最优目标预测器;(iv)对于线性模型,简单解析调整即可得到最优目标参数的一致估计。在合成和半合成数据上的实验验证了当假设成立时我们方法的有效性。最后,我们讨论了丰富的未来扩展方向。