AI-generated synthetic data, in addition to protecting the privacy of original data sets, allows users and data consumers to tailor data to their needs. This paper explores the creation of synthetic data that embodies Fairness by Design, focusing on the statistical parity fairness definition. By equalizing the learned target probability distributions of the synthetic data generator across sensitive attributes, a downstream model trained on such synthetic data provides fair predictions across all thresholds, that is, strong fair predictions even when inferring from biased, original data. This fairness adjustment can be either directly integrated into the sampling process of a synthetic generator or added as a post-processing step. The flexibility allows data consumers to create fair synthetic data and fine-tune the trade-off between accuracy and fairness without any previous assumptions on the data or re-training the synthetic data generator.
翻译:人工智能生成的合成数据,除保护原始数据集隐私外,还允许用户和数据消费者根据需求定制数据。本文探讨了体现"设计即公平"理念的合成数据生成方法,重点聚焦于统计均等性公平定义。通过均衡合成数据生成器在敏感属性上的学习目标概率分布,基于此类合成数据训练的下游模型可在所有决策阈值上提供公平预测,即使从有偏原始数据推断时也能实现强公平预测。该公平性调整可直接集成至合成数据生成器的采样过程,或作为后处理步骤添加。这种灵活性使数据消费者无需事先假设数据特征或重新训练生成器,即可创建公平合成数据并精细调节精度与公平性之间的平衡。