It is widely believed that given the same labeling budget, active learning (AL) algorithms like margin-based active learning achieve better predictive performance than passive learning (PL), albeit at a higher computational cost. Recent empirical evidence suggests that this added cost might be in vain, as margin-based AL can sometimes perform even worse than PL. While existing works offer different explanations in the low-dimensional regime, this paper shows that the underlying mechanism is entirely different in high dimensions: we prove for logistic regression that PL outperforms margin-based AL even for noiseless data and when using the Bayes optimal decision boundary for sampling. Insights from our proof indicate that this high-dimensional phenomenon is exacerbated when the separation between the classes is small. We corroborate this intuition with experiments on 20 high-dimensional datasets spanning a diverse range of applications, from finance and histology to chemistry and computer vision.
翻译:学界普遍认为,在相同标注预算下,基于边际的主动学习等主动学习算法虽计算成本更高,但仍能比被动学习取得更优的预测性能。最新实证证据表明,这种额外成本可能徒劳无益——基于边际的主动学习有时表现甚至劣于被动学习。现有研究在低维场景中提出了不同解释,但本文证明高维场景中的根本机制截然不同:我们通过逻辑回归证明,即便在无噪声数据且采用贝叶斯最优决策边界采样的情况下,被动学习仍优于基于边际的主动学习。证明过程中的洞见表明,当类别间分离度较小时,这种高维现象会进一步加剧。我们通过20个覆盖金融、组织学、化学与计算机视觉等多元领域的高维数据集实验,验证了这一直观结论。