We present a detailed study of top-$k$ classification, the task of predicting the $k$ most probable classes for an input, extending beyond single-class prediction. We demonstrate that several prevalent surrogate loss functions in multi-class classification, such as comp-sum and constrained losses, are supported by $H$-consistency bounds with respect to the top-$k$ loss. These bounds guarantee consistency in relation to the hypothesis set $H$, providing stronger guarantees than Bayes-consistency due to their non-asymptotic and hypothesis-set specific nature. To address the trade-off between accuracy and cardinality $k$, we further introduce cardinality-aware loss functions through instance-dependent cost-sensitive learning. For these functions, we derive cost-sensitive comp-sum and constrained surrogate losses, establishing their $H$-consistency bounds and Bayes-consistency. Minimizing these losses leads to new cardinality-aware algorithms for top-$k$ classification. We report the results of extensive experiments on CIFAR-100, ImageNet, CIFAR-10, and SVHN datasets demonstrating the effectiveness and benefit of these algorithms.
翻译:我们针对top-$k$分类(即为输入预测概率最高的$k$个类别,超越单类别预测任务)进行了详细研究。我们证明,多类别分类中的多种常用替代损失函数(如comp-sum损失和约束损失)均支持关于top-$k$损失的$H$一致性界。这些界保证了相对于假设集$H$的一致性,由于其非渐近性和假设集特异性,比贝叶斯一致性提供了更强的保证。为平衡准确率与基数$k$之间的权衡,我们进一步通过实例相关的代价敏感学习引入基数感知损失函数。针对这些函数,我们推导出代价敏感的comp-sum损失和约束损失替代函数,并建立其$H$一致性界与贝叶斯一致性。最小化这些损失可产生用于top-$k$分类的新型基数感知算法。我们在CIFAR-100、ImageNet、CIFAR-10和SVHN数据集上的大量实验结果展示了这些算法的有效性与优势。