Ensuring a neural network is not relying on protected attributes (e.g., race, sex, age) for predictions is crucial in advancing fair and trustworthy AI. While several promising methods for removing attribute bias in neural networks have been proposed, their limitations remain under-explored. In this work, we mathematically and empirically reveal an important limitation of attribute bias removal methods in presence of strong bias. Specifically, we derive a general non-vacuous information-theoretical upper bound on the performance of any attribute bias removal method in terms of the bias strength. We provide extensive experiments on synthetic, image, and census datasets to verify the theoretical bound and its consequences in practice. Our findings show that existing attribute bias removal methods are effective only when the inherent bias in the dataset is relatively weak, thus cautioning against the use of these methods in smaller datasets where strong attribute bias can occur, and advocating the need for methods that can overcome this limitation.
翻译:确保神经网络在预测时不依赖受保护属性(如种族、性别、年龄)对于推进公平可信的人工智能至关重要。尽管已有多种消除神经网络中属性偏差的方法被提出,但其局限性仍未得到充分探索。本研究从数学和实证角度揭示了属性偏差消除方法在强偏差存在下的重要局限性。具体而言,我们推导出任意属性偏差消除方法在偏差强度方面的通用非平凡信息论上界。通过在合成数据集、图像数据集和人口普查数据集上的大量实验,我们验证了该理论界及其在实际场景中的影响。研究结果表明,现有属性偏差消除方法仅在数据集固有偏差相对较弱时有效,因此需警惕在可能产生强属性偏差的小规模数据集中使用这些方法,并倡导开发能克服此局限性的新方法。