We propose an experimental method for measuring bias in face recognition systems. Existing methods to measure bias depend on benchmark datasets that are collected in the wild and annotated for protected (e.g., race, gender) and non-protected (e.g., pose, lighting) attributes. Such observational datasets only permit correlational conclusions, e.g., "Algorithm A's accuracy is different on female and male faces in dataset X.". By contrast, experimental methods manipulate attributes individually and thus permit causal conclusions, e.g., "Algorithm A's accuracy is affected by gender and skin color." Our method is based on generating synthetic faces using a neural face generator, where each attribute of interest is modified independently while leaving all other attributes constant. Human observers crucially provide the ground truth on perceptual identity similarity between synthetic image pairs. We validate our method quantitatively by evaluating race and gender biases of three research-grade face recognition models. Our synthetic pipeline reveals that for these algorithms, accuracy is lower for Black and East Asian population subgroups. Our method can also quantify how perceptual changes in attributes affect face identity distances reported by these models. Our large synthetic dataset, consisting of 48,000 synthetic face image pairs (10,200 unique synthetic faces) and 555,000 human annotations (individual attributes and pairwise identity comparisons) is available to researchers in this important area.
翻译:我们提出了一种用于测量面部识别系统偏差的实验方法。现有的偏差测量方法依赖于在自然环境中采集并标注受保护(如种族、性别)与非受保护(如姿态、光照)属性的基准数据集。此类观测性数据集仅能得出相关性结论,例如“算法A在数据集X中对女性与男性面部的准确率存在差异”。相比之下,实验方法通过独立操控属性可实现因果性结论,例如“算法A的准确率受性别与肤色影响”。本方法基于神经面部生成器合成人脸,在保持其他属性不变的前提下独立调整每个目标属性。人类观察者关键性地为合成图像对之间的感知身份相似性提供真值标注。我们通过评估三个研究级面部识别模型的种族与性别偏差,对方法进行了定量验证。合成管道表明,这些算法对黑人与东亚人群子组的准确率较低。本方法还可量化属性感知变化对这些模型报告的面部身份距离的影响。我们的大型合成数据集包含48,000对合成人脸图像(10,200张独立合成人脸)及555,000条人类标注(个体属性与成对身份比较),现可供该重要领域的研究人员使用。