The use of machine learning models in decision support systems with high societal impact raised concerns about unfair (disparate) results for different groups of people. When evaluating such unfair decisions, one generally relies on predefined groups that are determined by a set of features that are considered sensitive. However, such an approach is subjective and does not guarantee that these features are the only ones to be considered as sensitive nor that they entail unfair (disparate) outcomes. In this paper, we propose a preprocessing step to address the task of automatically recognizing sensitive features that does not require a trained model to verify unfair results. Our proposal is based on the Hilber-Schmidt independence criterion, which measures the statistical dependence of variable distributions. We hypothesize that if the dependence between the label vector and a candidate is high for a sensitive feature, then the information provided by this feature will entail disparate performance measures between groups. Our empirical results attest our hypothesis and show that several features considered as sensitive in the literature do not necessarily entail disparate (unfair) results.
翻译:机器学习模型广泛应用于具有重大社会影响的决策支持系统,引发了不同人群间可能产生不公平(差异性)结果的担忧。在评估此类不公决策时,通常依赖于由一组被视为敏感的特征所定义的预定群体。然而,这种方法具有主观性,既无法确保这些特征是唯一需要被视为敏感的特征,也无法保证它们必然导致不公平(差异性)结果。本文提出一种预处理步骤,用于自动识别敏感特征,且无需借助已训练模型来验证不公结果。该方案基于希尔伯特-施密特独立性准则,该准则用于衡量变量分布间的统计依赖性。我们假设:若标签向量与某个候选特征之间的依赖程度对敏感特征而言较高,则该特征所提供的信息将导致群体间的性能指标差异。实证结果验证了该假设,并表明文献中多个被视为敏感的特征未必必然导致差异性(不公平)结果。