Most group fairness notions detect unethical biases by computing statistical parity metrics on a model's output. However, this approach suffers from several shortcomings, such as philosophical disagreement, mutual incompatibility, and lack of interpretability. These shortcomings have spurred the research on complementary bias detection methods that offer additional transparency into the sources of discrimination and are agnostic towards an a priori decision on the definition of fairness and choice of protected features. A recent proposal in this direction is LUCID (Locating Unfairness through Canonical Inverse Design), where canonical sets are generated by performing gradient descent on the input space, revealing a model's desired input given a preferred output. This information about the model's mechanisms, i.e., which feature values are essential to obtain specific outputs, allows exposing potential unethical biases in its internal logic. Here, we present LUCID-GAN, which generates canonical inputs via a conditional generative model instead of gradient-based inverse design. LUCID-GAN has several benefits, including that it applies to non-differentiable models, ensures that canonical sets consist of realistic inputs, and allows to assess proxy and intersectional discrimination. We empirically evaluate LUCID-GAN on the UCI Adult and COMPAS data sets and show that it allows for detecting unethical biases in black-box models without requiring access to the training data.
翻译:摘要:大多数群体公平性概念通过计算模型输出上的统计平价指标来检测不道德偏见。然而,这种方法存在若干缺陷,如哲学争议、相互不兼容以及缺乏可解释性。这些缺陷催生了对互补性偏见检测方法的研究,这些方法能提供关于歧视来源的额外透明度,且不受先验确定的公平性定义和受保护特征选择的影响。近期该方向的一项提议是LUCID(通过规范逆设计定位不公平性),该方法通过在输入空间上执行梯度下降生成规范集,揭示模型在给定首选输出下的期望输入。关于模型机制的信息(即哪些特征值对获得特定输出至关重要)有助于暴露其内部逻辑中潜在的不道德偏见。本文提出LUCID-GAN,该方法通过条件生成模型而非基于梯度的逆设计生成规范输入。LUCID-GAN具有多项优势,包括适用于非可微模型、确保规范集由真实输入构成,以及支持评估代理歧视和交叉歧视。我们在UCI Adult和COMPAS数据集上对LUCID-GAN进行实证评估,结果表明该方法无需访问训练数据即可检测黑盒模型中的不道德偏见。