We consider the problem of learning from data corrupted by underrepresentation bias, where positive examples are filtered from the data at different, unknown rates for a fixed number of sensitive groups. We show that with a small amount of unbiased data, we can efficiently estimate the group-wise drop-out rates, even in settings where intersectional group membership makes learning each intersectional rate computationally infeasible. Using these estimates, we construct a reweighting scheme that allows us to approximate the loss of any hypothesis on the true distribution, even if we only observe the empirical error on a biased sample. From this, we present an algorithm encapsulating this learning and reweighting process along with a thorough empirical investigation. Finally, we define a bespoke notion of PAC learnability for the underrepresentation and intersectional bias setting and show that our algorithm permits efficient learning for model classes of finite VC dimension.
翻译:我们研究从受代表性不足偏差污染的数据中学习的问题,其中正例以不同且未知的速率从固定数量的敏感群体数据中被过滤。我们证明,利用少量无偏数据,即使在学习每个交叉群体成员对应的速率计算不可行的情况下,仍能有效估计各群体的数据丢弃率。基于这些估计,我们构建了一种重加权方案,使得即使在仅观测到有偏样本上的经验误差时,也能近似计算任意假设在真实分布上的损失。基于此,我们提出了一种整合该学习与重加权过程的算法,并进行了详尽的实证研究。最后,我们为代表性不足与交叉性偏差场景定义了定制化的PAC可学习性概念,并证明我们的算法允许对有限VC维度的模型类实现高效学习。