Despite impressive advances in object-recognition, deep learning systems' performance degrades significantly across geographies and lower income levels raising pressing concerns of inequity. Addressing such performance gaps remains a challenge, as little is understood about why performance degrades across incomes or geographies. We take a step in this direction by annotating images from Dollar Street, a popular benchmark of geographically and economically diverse images, labeling each image with factors such as color, shape, and background. These annotations unlock a new granular view into how objects differ across incomes and regions. We then use these object differences to pinpoint model vulnerabilities across incomes and regions. We study a range of modern vision models, finding that performance disparities are most associated with differences in texture, occlusion, and images with darker lighting. We illustrate how insights from our factor labels can surface mitigations to improve models' performance disparities. As an example, we show that mitigating a model's vulnerability to texture can improve performance on the lower income level. We release all the factor annotations along with an interactive dashboard to facilitate research into more equitable vision systems.
翻译:尽管物体识别技术取得了显著进步,但深度学习系统的性能在不同地理区域和较低收入水平下显著下降,这引发了关于不公平性的紧迫担忧。解决此类性能差距仍是一项挑战,因为人们对性能为何随收入或地理区域下降知之甚少。我们通过为Dollar Street(一个广泛使用的、涵盖地理和经济多样性图像的数据集)中的图像添加注释,标注每张图像的颜色、形状和背景等因素,向这一方向迈出了一步。这些注释揭示了物体在不同收入和地区间差异的新视角。随后,我们利用这些物体差异来定位模型在不同收入和地区的脆弱性。我们研究了一系列现代视觉模型,发现性能差异最显著地与纹理、遮挡和较暗光照图像的差异相关。我们展示了如何通过因素标签中的洞察来提出缓解措施,以改善模型的性能差距。例如,我们表明缓解模型对纹理的脆弱性可以提升其在低收入水平下的性能。我们发布了所有因素注释以及一个交互式仪表板,以促进对更公平视觉系统的研究。