Deep neural networks (DNNs) are suffering from ethical issues such as individual discrimination. In response, extensive NN repair techniques have been developed to adjust models and mitigate such undesired behaviors. However, existing fairness repair methods are typically data-centric, which often lack provable guarantees and generalization to unseen samples. To overcome these limitations, we propose ProF, a novel fairness repair framework with provable guarantees. The key intuition of ProF is to leverage interval bound propagation (a widely used NN verification technique) to soundly capture model outputs over the whole set $S(\mathbf{x})$ around a biased sample $\mathbf{x}$. The derived bounds are utilized to guide fairness repair which encourages the model to produce consistent outputs on $S(\mathbf{x})$. Specifically, we integrate fairness constraints and model modifications into a unified constraint-solving formulation, which can be transformed to a Mixed-Integer Linear Programming (MILP) problem solvable by off-the-shelf solvers. The solution to the MILP problem effectively induces a repaired model with guaranteed fairness over the whole set $S(\mathbf{x})$. We evaluate ProF on four widely used benchmark datasets and demonstrate that it achieves provable fairness repair, with generalization of up to 95.93\% on full datasets and 93.16\% on the entire input space. Notably, ProF can be easily configured to support multiple sensitive attributes and more practical fairness definitions, while providing provable repair guarantees and delivering around 90\% fairness improvement. Our code is available at https://github.com/nninjn/ProF.


翻译:深度神经网络(DNN)正面临诸如个体歧视等伦理问题。为此,研究人员开发了大量神经网络修复技术,以调整模型并缓解此类不当行为。然而,现有的公平性修复方法通常以数据为中心,往往缺乏可证明的保障以及对未见样本的泛化能力。为克服这些局限,我们提出ProF——一种具有可证明保障的新型公平性修复框架。ProF的核心思想是利用区间边界传播(一种广泛使用的神经网络验证技术)来可靠地捕捉有偏样本$\mathbf{x}$周围整个集合$S(\mathbf{x})$上的模型输出。所推导出的边界被用于指导公平性修复,促使模型在$S(\mathbf{x})$上产生一致的输出。具体而言,我们将公平性约束和模型修改整合为一个统一的约束求解公式,该公式可转化为可由现成求解器求解的混合整数线性规划(MILP)问题。MILP问题的解能有效诱导出修复后的模型,且该模型在整体集合$S(\mathbf{x})$上具有有保障的公平性。我们在四个广泛使用的基准数据集上评估了ProF,结果表明它实现了可证明的公平性修复,在完整数据集上的泛化能力高达95.93%,在整个输入空间上高达93.16%。值得注意的是,ProF可轻松配置以支持多个敏感属性和更实用的公平性定义,同时提供可证明的修复保障,并实现约90%的公平性提升。我们的代码可在https://github.com/nninjn/ProF获取。

0
下载
关闭预览

相关内容

基于深度神经网络的图像缺损修复方法综述
专知会员服务
26+阅读 · 2021年12月18日
专知会员服务
171+阅读 · 2020年8月26日
最新《可解释深度学习XDL》2020研究进展综述大全,54页pdf
深度学习模型可解释性的研究进展
专知
26+阅读 · 2020年8月1日
【GNN】深度学习之上,图神经网络(GNN )崛起
产业智能官
16+阅读 · 2019年8月15日
用深度学习揭示数据的因果关系
专知
28+阅读 · 2019年5月18日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
VIP会员
最新内容
美海军陆战队将三型无人机整合入统一战场网络
专知会员服务
2+阅读 · 今天9:39
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
3+阅读 · 今天9:12
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 今天9:08
美陆军-工业界协同推进反无人机系统技术发展
专知会员服务
1+阅读 · 今天8:46
《跨域指挥背景下的领导力发展》最新报告
专知会员服务
1+阅读 · 今天8:40
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员