Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques have been designed to replace missing data with stand in values. The various approaches have implications for calculating clinical scores, model building and model testing. The work showcased here offers a novel means for categorical imputation based on item response theory (IRT) and compares it against several methodologies currently used in the machine learning field including k-nearest neighbors (kNN), multiple imputed chained equations (MICE) and Amazon Web Services (AWS) deep learning method, Datawig. Analyses comparing these techniques were performed on three different datasets that represented ordinal, nominal and binary categories. The data were modified so that they also varied on both the proportion of data missing and the systematization of the missing data. Two different assessments of performance were conducted: accuracy in reproducing the missing values, and predictive performance using the imputed data. Results demonstrated that the new method, Item Response Theory for Categorical Imputation (IRTCI), fared quite well compared to currently used methods, outperforming several of them in many conditions. Given the theoretical basis for the new approach, and the unique generation of probabilistic terms for determining category belonging for missing cells, IRTCI offers a viable alternative to current approaches.
翻译:大多数数据集存在部分或完全缺失值的问题,这对可用于测试数据的模型以及从数据中得出的统计推断造成了下游限制。目前已设计出多种插补技术,用替代值替换缺失数据。不同方法对临床评分计算、模型构建和模型测试均有影响。本研究提出了一种基于项目反应理论(IRT)的分类数据插补新方法,并将其与当前机器学习领域的几种常用方法进行了比较,包括k近邻(kNN)、多重插补链式方程(MICE)以及亚马逊云服务(AWS)深度学习方法Datawig。为进行比较分析,我们在三个分别代表有序、名义和二元分类的数据集上进行了实验。通过修改数据,使其在缺失数据比例和缺失数据系统性两方面均存在差异。我们采用两种不同的性能评估方式:缺失值还原的准确率,以及使用插补数据后的预测性能。结果表明,新方法——基于项目反应理论的分类插补(IRTCI)——与现有方法相比表现良好,在许多条件下优于其中若干方法。鉴于该新方法具有理论基础,且能独特地生成概率项以确定缺失单元格的类别归属,IRTCI为现有方法提供了一种可行的替代方案。