Missing data is a ubiquitous challenge in data analysis, often leading to biased and inaccurate results. Traditional imputation methods usually assume that the missingness mechanism is missing-at-random (MAR), where the missingness is independent of the missing values themselves. This assumption is frequently violated in real-world scenarios, prompted by recent advances in imputation methods using deep learning to address this challenge. However, these methods neglect the crucial issue of nonparametric identifiability in missing-not-at-random (MNAR) data, which can lead to biased and unreliable results. This paper seeks to bridge this gap by proposing a novel framework based on deep latent variable models for {MNAR data}. Building on the assumption of conditional no self-censoring {given} latent variables, we establish the identifiability of the data distribution. This crucial theoretical result guarantees the feasibility of our approach. To effectively estimate unknown parameters, we develop an efficient algorithm utilizing importance-weighted autoencoders. We demonstrate, both theoretically and empirically, that our estimation process accurately recovers the ground-truth joint distribution under specific regularity conditions. Extensive simulation studies and real-world data experiments showcase the advantages of our proposed method compared to various classical and state-of-the-art approaches to missing data imputation.
翻译:缺失数据是数据分析中普遍存在的挑战,常导致有偏且不精确的结果。传统插补方法通常假设缺失机制为随机缺失(MAR),即缺失与否与缺失值本身无关。然而,这一假设在现实场景中常被违反,近期基于深度学习的插补方法进展旨在应对这一挑战。但这些方法忽略了非随机缺失(MNAR)数据中非参数可辨识性的关键问题,可能导致有偏且不可靠的结果。本文旨在弥合这一缺口,提出一种基于深度潜变量模型的新型框架以处理MNAR数据。基于潜在变量条件下无自我删失的假设,我们确立了数据分布的可辨识性。这一关键理论结果保证了我们方法的可行性。为有效估计未知参数,我们开发了一种利用重要性加权自编码器的高效算法。我们从理论与实证两方面证明,在特定正则条件下,我们的估计过程能准确恢复真实联合分布。广泛的模拟研究与真实数据实验展示了我们方法相较于多种经典及前沿缺失数据插补方法的优势。