Cost-sensitive learning is a common type of machine learning problem where different errors of prediction incur different costs. In this paper, we design a generic nonparametric active learning algorithm for cost-sensitive classification. Based on the construction of confidence bounds for the expected prediction cost functions of each label, our algorithm sequentially selects the most informative vector points. Then it interacts with them by only querying the costs of prediction that could be the smallest. We prove that our algorithm attains optimal rate of convergence in terms of the number of interactions with the feature vector space. Furthermore, in terms of a general version of Tsybakov's noise assumption, the gain over the corresponding passive learning is explicitly characterized by the probability-mass of the boundary decision. Additionally, we prove the near-optimality of obtained upper bounds by providing matching (up to logarithmic factor) lower bounds.
翻译:代价敏感学习是机器学习中的常见问题类型,其中不同的预测错误会导致不同的代价。本文针对代价敏感分类问题,设计了一种通用的非参数主动学习算法。该算法基于为每个标签的期望预测代价函数构建置信区间,依次选择信息量最大的特征向量点,并通过仅查询可能产生最小预测代价的样本来与之交互。我们证明,该算法在与特征向量空间交互次数方面达到了最优收敛速率。此外,在泰巴科夫噪声假设的一般版本下,相较于相应的被动学习,其性能增益由决策边界的概率质量显式表征。进一步地,我们通过提供匹配(至多相差对数因子)的下界,证明了所得上界的近最优性。