Although several fairness definitions and bias mitigation techniques exist in the literature, all existing solutions evaluate fairness of Machine Learning (ML) systems after the training stage. In this paper, we take the first steps towards evaluating a more holistic approach by testing for fairness both before and after model training. We evaluate the effectiveness of the proposed approach and position it within the ML development lifecycle, using an empirical analysis of the relationship between model dependent and independent fairness metrics. The study uses 2 fairness metrics, 4 ML algorithms, 5 real-world datasets and 1600 fairness evaluation cycles. We find a linear relationship between data and model fairness metrics when the distribution and the size of the training data changes. Our results indicate that testing for fairness prior to training can be a ``cheap'' and effective means of catching a biased data collection process early; detecting data drifts in production systems and minimising execution of full training cycles thus reducing development time and costs.
翻译:尽管现有文献中已存在多种公平性定义和偏差缓解技术,但所有现有解决方案均是在训练阶段之后评估机器学习系统的公平性。本文首次探索更全面的评估方法,通过在模型训练前后两个阶段测试公平性。我们通过实证分析模型相关指标与模型无关公平性指标之间的关系,评估所提方法的有效性,并将其定位在机器学习开发生命周期中。本研究采用2项公平性指标、4种机器学习算法、5个真实世界数据集及1600次公平性评估循环。研究发现,当训练数据的分布和规模发生变化时,数据公平性指标与模型公平性指标之间存在线性关系。研究结果表明,训练前进行公平性测试是一种"低成本"的有效手段,能够早期发现存在偏差的数据收集过程、检测生产系统中的数据漂移,并减少完整训练周期的执行次数,从而降低开发时间和成本。