In machine learning, training data often capture the behaviour of multiple subgroups of some underlying human population. When the nature of training data for subgroups are not controlled carefully, under-representation bias arises. To counter this effect we introduce two natural notions of subgroup fairness and instantaneous fairness to address such under-representation bias in time-series forecasting problems. Here we show globally convergent methods for the fairness-constrained learning problems using hierarchies of convexifications of non-commutative polynomial optimisation problems. Our empirical results on a biased data set motivated by insurance applications and the well-known COMPAS data set demonstrate the efficacy of our methods. We also show that by exploiting sparsity in the convexifications, we can reduce the run time of our methods considerably.
翻译:在机器学习中,训练数据常包含某一人群中多个子群体的行为模式。当各子群体训练数据的性质未得到谨慎控制时,便会产生代表性偏差。为解决这一问题,我们在时间序列预测问题中引入了两种自然的子群体公平性与瞬时公平性概念。本文展示了利用非交换多项式优化问题的凸化分层结构,实现公平性约束学习问题的全局收敛方法。基于保险应用场景的偏倚数据集与著名的COMPAS数据集的实证结果,验证了所提方法的有效性。进一步研究表明,通过利用凸化过程中的稀疏性,可显著降低方法的运行时间。