Data imbalance in training data often leads to biased predictions from trained models, which in turn causes ethical and social issues. A straightforward solution is to carefully curate training data, but given the enormous scale of modern neural networks, this is prohibitively labor-intensive and thus impractical. Inspired by recent developments in generative models, this paper explores the potential of synthetic data to address the data imbalance problem. To be specific, our method, dubbed SYNAuG, leverages synthetic data to equalize the unbalanced distribution of training data. Our experiments demonstrate that, although a domain gap between real and synthetic data exists, training with SYNAuG followed by fine-tuning with a few real samples allows to achieve impressive performance on diverse tasks with different data imbalance issues, surpassing existing task-specific methods for the same purpose.
翻译:训练数据中的不均衡常导致模型预测产生偏差,进而引发伦理与社会问题。一个直接的解决方案是精心筛选训练数据,但鉴于现代神经网络的庞大规模,这一方法劳动成本过高且不切实际。受生成模型最新进展的启发,本文探索了合成数据在解决数据不平衡问题中的潜力。具体而言,我们提出的方法SYNAuG通过合成数据来均衡训练数据的不平衡分布。实验表明,尽管真实数据与合成数据之间存在领域差异,但采用SYNAuG训练后,仅需少量真实样本进行微调,即可在多种存在不同数据不平衡问题的任务中取得显著性能,超越了现有针对相同目的设计的特定任务方法。