Nowadays an ever-growing concerning phenomenon, the emergence of algorithmic biases that can lead to unfair models, emerges. Several debiasing approaches have been proposed in the realm of deep learning, employing more or less sophisticated approaches to discourage these models from massively employing these biases. However, a question emerges: is this extra complexity really necessary? Is a vanilla-trained model already embodying some ``unbiased sub-networks'' that can be used in isolation and propose a solution without relying on the algorithmic biases? In this work, we show that such a sub-network typically exists, and can be extracted from a vanilla-trained model without requiring additional training. We further validate that such specific architecture is incapable of learning a specific bias, suggesting that there are possible architectural countermeasures to the problem of biases in deep neural networks.
翻译:如今,一个日益令人担忧的现象——可能导致不公平模型的算法偏差——开始浮现。在深度学习领域,已有多种去偏方法被提出,它们采用或简或繁的策略来抑制模型大规模利用这些偏差。然而,一个问题随之而来:这种额外的复杂性真的必要吗?一个经过普通训练(vanilla-trained)的模型是否已经包含某些“无偏子网络”,这些子网络可以单独使用,从而提出不依赖算法偏差的解决方案?在本工作中,我们证明这类子网络通常存在,并且可以从普通训练模型中提取,无需额外训练。我们进一步验证,这种特定架构无法学习特定偏差,这表明在深度神经网络中可能存在针对偏差问题的架构性对策。