AI systems have been known to amplify biases in real-world data. Explanations may help human-AI teams address these biases for fairer decision-making. Typically, explanations focus on salient input features. If a model is biased against some protected group, explanations may include features that demonstrate this bias, but when biases are realized through proxy features, the relationship between this proxy feature and the protected one may be less clear to a human. In this work, we study the effect of the presence of protected and proxy features on participants' perception of model fairness and their ability to improve demographic parity over an AI alone. Further, we examine how different treatments -- explanations, model bias disclosure and proxy correlation disclosure -- affect fairness perception and parity. We find that explanations help people detect direct but not indirect biases. Additionally, regardless of bias type, explanations tend to increase agreement with model biases. Disclosures can help mitigate this effect for indirect biases, improving both unfairness recognition and decision-making fairness. We hope that our findings can help guide further research into advancing explanations in support of fair human-AI decision-making.
翻译:人工智能系统已被证实会放大现实世界数据中的偏见。解释可能有助于人类-人工智能协作团队解决这些偏见,从而做出更公平的决策。通常,解释聚焦于显著的输入特征。若模型对某些受保护群体存在偏见,解释可能包含体现此偏见的特征;但当偏见通过代理特征实现时,该代理特征与受保护特征间的关系对人类而言可能较不清晰。本研究探讨了受保护特征与代理特征的存在如何影响参与者对模型公平性的感知,以及他们提升较之纯人工智能系统的群体均等性的能力。进一步地,我们检验了不同处理方式——解释、模型偏见披露及代理相关性披露——如何影响公平性感知与均等性。我们发现,解释有助于人们检测直接偏见,但无法检测间接偏见。此外,无论偏见类型如何,解释往往增加对模型偏见的认同度。披露措施可缓解间接偏见的这种影响,既改善了对不公平的识别,也提升了决策公平性。我们期望这些发现能引导后续研究,推动解释方法的发展以支持公平的人类-人工智能决策。