Missing data imputation is a fundamental problem in data analysis, and many studies have been conducted to improve its performance by exploring model structures and learning procedures. However, data augmentation, as a simple yet effective method, has not received enough attention in this area. In this paper, we propose a novel data augmentation method called Missingness Augmentation (MisA) for generative imputation models. Our approach dynamically produces incomplete samples at each epoch by utilizing the generator's output, constraining the augmented samples using a simple reconstruction loss, and combining this loss with the original loss to form the final optimization objective. As a general augmentation technique, MisA can be easily integrated into generative imputation frameworks, providing a simple yet effective way to enhance their performance. Experimental results demonstrate that MisA significantly improves the performance of many recently proposed generative imputation models on a variety of tabular and image datasets. The code is available at \url{https://github.com/WYu-Feng/Missingness-Augmentation}.


翻译:缺失数据插补是数据分析中的基本问题,许多研究通过探索模型结构和学习流程来提升其性能。然而,数据增强作为一种简单而有效的方法,在该领域尚未得到充分关注。本文提出一种名为缺失增强(MisA)的新型数据增强方法,专门用于生成式插补模型。该方法在每个训练周期利用生成器的输出动态生成不完整样本,通过简单的重构损失约束增强后的样本,并将该损失与原始损失结合构成最终优化目标。作为一种通用增强技术,MisA可轻松集成到生成式插补框架中,提供一种简单有效的性能提升方式。实验结果表明,MisA在多种表格和图像数据集上显著提升了近期提出的多个生成式插补模型的性能。代码开源在\url{https://github.com/WYu-Feng/Missingness-Augmentation}。

0
下载
关闭预览

相关内容

【NeurIPS2022】隐空间变换解决GAN生成分布的非连续性问题
专知会员服务
26+阅读 · 2022年11月30日
专知会员服务
63+阅读 · 2020年3月4日
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
深度自进化聚类:Deep Self-Evolution Clustering
我爱读PAMI
15+阅读 · 2019年4月13日
逆强化学习-学习人先验的动机
CreateAMind
16+阅读 · 2019年1月18日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
vae 相关论文 表示学习 1
CreateAMind
12+阅读 · 2018年9月6日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Arxiv
0+阅读 · 2023年5月24日
Arxiv
0+阅读 · 2023年5月23日
Arxiv
14+阅读 · 2022年8月25日
Arxiv
10+阅读 · 2021年3月30日
Feature Denoising for Improving Adversarial Robustness
Arxiv
15+阅读 · 2018年12月9日
VIP会员
最新内容
《多域冲突比较支持模型》60页
专知会员服务
4+阅读 · 今天4:35
面向2027年及未来的海军情报改革
专知会员服务
4+阅读 · 8月5日
相关VIP内容
【NeurIPS2022】隐空间变换解决GAN生成分布的非连续性问题
专知会员服务
26+阅读 · 2022年11月30日
专知会员服务
63+阅读 · 2020年3月4日
相关基金
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员