When dealing with imbalanced classification data, reweighting the loss function is a standard procedure allowing to equilibrate between the true positive and true negative rates within the risk measure. Despite significant theoretical work in this area, existing results do not adequately address a main challenge within the imbalanced classification framework, which is the negligible size of one class in relation to the full sample size and the need to rescale the risk function by a probability tending to zero. To address this gap, we present two novel contributions in the setting where the rare class probability approaches zero: (1) a non asymptotic fast rate probability bound for constrained balanced empirical risk minimization, and (2) a consistent upper bound for balanced nearest neighbors estimates. Our findings provide a clearer understanding of the benefits of class-weighting in realistic settings, opening new avenues for further research in this field.
翻译:在处理不平衡分类数据时,对损失函数进行重加权是一种标准做法,它允许在风险度量中平衡真正例率和真负例率。尽管该领域已有大量理论工作,但现有结果并未充分解决不平衡分类框架内的一个主要挑战——即某一类样本相对于总样本量而言规模可忽略不计,且需要以趋近于零的概率重新缩放风险函数。为弥补这一空白,我们在稀有类概率趋近于零的设定下提出两项新颖贡献:(1)约束平衡经验风险最小化的非渐近快速率概率界,以及(2)平衡最近邻估计的一致上界。我们的研究结果更清晰地揭示了在实际场景中类别加权的好处,为该领域的进一步研究开辟了新途径。