Machine learning systems require representations of the real world for training and testing - they require data, and lots of it. Collecting data at scale has logistical and ethical challenges, and synthetic data promises a solution to these challenges. Instead of needing to collect photos of real people's faces to train a facial recognition system, a model creator could create and use photo-realistic, synthetic faces. The comparative ease of generating this synthetic data rather than relying on collecting data has made it a common practice. We present two key risks of using synthetic data in model development. First, we detail the high risk of false confidence when using synthetic data to increase dataset diversity and representation. We base this in the examination of a real world use-case of synthetic data, where synthetic datasets were generated for an evaluation of facial recognition technology. Second, we examine how using synthetic data risks circumventing consent for data usage. We illustrate this by considering the importance of consent to the U.S. Federal Trade Commission's regulation of data collection and affected models. Finally, we discuss how these two risks exemplify how synthetic data complicates existing governance and ethical practice; by decoupling data from those it impacts, synthetic data is prone to consolidating power away those most impacted by algorithmically-mediated harm.
翻译:机器学习系统需要真实世界的表征来进行训练和测试——它们需要数据,而且是大量数据。大规模收集数据面临后勤和伦理方面的挑战,而合成数据有望解决这些挑战。模型创建者无需收集真实人脸照片来训练面部识别系统,而是可以创建并使用逼真的合成人脸。生成这种合成数据相对容易,且无需依赖数据收集,这使其成为一种常见做法。我们提出了在模型开发中使用合成数据的两个关键风险。首先,我们详细阐述了在使用合成数据提高数据集多样性和代表性时产生虚假自信的高风险。这一论点基于对合成数据真实用例的考察——在该用例中,为评估面部识别技术生成了合成数据集。其次,我们探讨了使用合成数据如何可能规避数据使用同意。我们通过分析同意对美国联邦贸易委员会数据收集及受影响模型监管的重要性来阐述这一点。最后,我们讨论了这两个风险如何体现合成数据使现有治理和伦理实践复杂化:通过将数据与其影响对象脱钩,合成数据容易将权力集中到远离算法中介危害最深者的一方。