The potential harms of the under-representation of minorities in training data, particularly in multi-modal settings, is a well-recognized concern. While there has been extensive effort in detecting such under-representation, resolution has remained a challenge. With recent advancements in generative AI, large language models and foundation models have emerged as versatile tools across various domains. In this paper, we propose Chameleon, a system that efficiently utilizes these tools to augment a data set with a minimal addition of synthetically generated tuples, in order to enhance the coverage of the under-represented groups. Our system follows a rejection sampling approach to ensure the generated tuples have a high quality and follow the underlying distribution. In order to minimize the rejection chance of the generated tuples, we propose multiple strategies for providing a guide for the foundation model. Our experiment results, in addition to confirming the efficiency of our proposed algorithms, illustrate the effectiveness of our approach, as the unfairness of the model in a downstream task significantly dropped after data repair using Chameleon.
翻译:训练数据中少数群体代表性不足的潜在危害,特别是在多模态场景下,已成为广受关注的问题。尽管已有大量研究致力于检测此类代表性不足现象,但解决该问题仍面临挑战。随着生成式人工智能的最新进展,大语言模型和基础模型已成为跨越多个领域的通用工具。本文提出Chameleon系统,该系统高效利用这些工具,通过最小化合成生成元组的补充来增强数据集,从而提升对代表性不足群体的覆盖。我们的系统采用拒绝采样方法,确保生成的元组具有高质量且遵循底层数据分布。为最小化生成元组的拒绝概率,我们提出了多种为基础模型提供引导的策略。实验结果不仅验证了所提算法的效率,更证明了该方法的有效性:经Chameleon修复数据后,下游任务中的模型不公平性显著降低。