Seven years ago, researchers proposed a postprocessing method to equalize the error rates of a model across different demographic groups. The work launched hundreds of papers purporting to improve over the postprocessing baseline. We empirically evaluate these claims through thousands of model evaluations on several tabular datasets. We find that the fairness-accuracy Pareto frontier achieved by postprocessing contains all other methods we were feasibly able to evaluate. In doing so, we address two common methodological errors that have confounded previous observations. One relates to the comparison of methods with different unconstrained base models. The other concerns methods achieving different levels of constraint relaxation. At the heart of our study is a simple idea we call unprocessing that roughly corresponds to the inverse of postprocessing. Unprocessing allows for a direct comparison of methods using different underlying models and levels of relaxation.
翻译:七年前,研究者提出了一种后处理方法,旨在均衡模型在不同人口群体间的错误率。该工作催生了数百篇声称优于后处理基线的论文。我们通过对多个表格数据集进行数千次模型评估,实证检验了这些声明。研究发现,后处理方法所实现的公平-准确率帕累托前沿包含了所有其他我们可实际评估的方法。在此过程中,我们纠正了先前观察中存在的两个常见方法论错误:其一涉及不同无约束基础模型间的比较方法;其二涉及不同约束松弛程度的技术方案。本研究的核心是一种名为"解构"的简单理念,其大致对应后处理的逆向操作。解构方法能够直接比较使用不同底层模型与松弛程度的各类技术方案。