The quintessential learning algorithm of empirical risk minimization (ERM) is known to fail in various settings for which uniform convergence does not characterize learning. It is therefore unsurprising that the practice of machine learning is rife with considerably richer algorithmic techniques for successfully controlling model capacity. Nevertheless, no such technique or principle has broken away from the pack to characterize optimal learning in these more general settings. The purpose of this work is to characterize the role of regularization in perhaps the simplest setting for which ERM fails: multiclass learning with arbitrary label sets. Using one-inclusion graphs (OIGs), we exhibit optimal learning algorithms that dovetail with tried-and-true algorithmic principles: Occam's Razor as embodied by structural risk minimization (SRM), the principle of maximum entropy, and Bayesian reasoning. Most notably, we introduce an optimal learner which relaxes structural risk minimization on two dimensions: it allows the regularization function to be "local" to datapoints, and uses an unsupervised learning stage to learn this regularizer at the outset. We justify these relaxations by showing that they are necessary: removing either dimension fails to yield a near-optimal learner. We also extract from OIGs a combinatorial sequence we term the Hall complexity, which is the first to characterize a problem's transductive error rate exactly. Lastly, we introduce a generalization of OIGs and the transductive learning setting to the agnostic case, where we show that optimal orientations of Hamming graphs -- judged using nodes' outdegrees minus a system of node-dependent credits -- characterize optimal learners exactly. We demonstrate that an agnostic version of the Hall complexity again characterizes error rates exactly, and exhibit an optimal learner using maximum entropy programs.
翻译:经验风险最小化(ERM)这一经典学习算法在多种均匀收敛无法描述学习过程的场景中已被证明会失效。因此,机器学习实践中广泛采用更为丰富的算法技术来成功控制模型容量也就不足为奇。然而,在这些更一般的场景中,尚未有任何技术或原理能够脱颖而出以刻画最优学习的特征。本研究旨在通过正则化的视角,在ERM失效的最简单场景——任意标签集的多类学习——中揭示其作用机制。利用单包含图(OIGs),我们展示了一系列与成熟算法原则相契合的最优学习算法:体现奥卡姆剃刀原则的结构风险最小化(SRM)、最大熵原理以及贝叶斯推理。尤为重要的是,我们提出了一种最优学习器,它在两个维度上放宽了结构风险最小化的限制:允许正则化函数对数据点具有“局部”特性,并利用无监督学习阶段在初始阶段学习该正则化器。我们通过证明这两个维度缺一不可来论证这种放宽的必要性:移除任一维度均无法得到接近最优的学习器。此外,我们从OIGs中提取出一种称为霍尔复杂度的组合序列,这是首个能精确刻画问题转导错误率的指标。最后,我们将OIGs与转导学习设置推广到不可知场景,证明了以节点出度减去节点相关信用系统为评价标准的汉明图最优定向能精确刻画最优学习器。我们验证了不可知版本的霍尔复杂度同样能精确描述错误率,并展示了基于最大熵规划的最优学习器。