The measure of a machine learning algorithm is the difficulty of the tasks it can perform, and sufficiently difficult tasks are critical drivers of strong machine learning models. However, quantifying the generalization difficulty of machine learning benchmarks has remained challenging. We propose what is to our knowledge the first model-agnostic measure of the inherent generalization difficulty of tasks. Our inductive bias complexity measure quantifies the total information required to generalize well on a task minus the information provided by the data. It does so by measuring the fractional volume occupied by hypotheses that generalize on a task given that they fit the training data. It scales exponentially with the intrinsic dimensionality of the space over which the model must generalize but only polynomially in resolution per dimension, showing that tasks which require generalizing over many dimensions are drastically more difficult than tasks involving more detail in fewer dimensions. Our measure can be applied to compute and compare supervised learning, reinforcement learning and meta-learning generalization difficulties against each other. We show that applied empirically, it formally quantifies intuitively expected trends, e.g. that in terms of required inductive bias, MNIST < CIFAR10 < Imagenet and fully observable Markov decision processes (MDPs) < partially observable MDPs. Further, we show that classification of complex images < few-shot meta-learning with simple images. Our measure provides a quantitative metric to guide the construction of more complex tasks requiring greater inductive bias, and thereby encourages the development of more sophisticated architectures and learning algorithms with more powerful generalization capabilities.
翻译:衡量机器学习算法能力的关键在于其可执行任务的难度,而足够困难的任务是推动强机器学习模型发展的核心驱动力。然而,量化机器学习基准(benchmark)的泛化难度仍具挑战性。本文提出了一种据我们所知首个与模型无关的任务固有泛化难度度量方法。我们的归纳偏置复杂度度量(inductive bias complexity measure)量化了在任务上实现良好泛化所需的总信息量减去数据已提供的信息量。该度量通过计算在拟合训练数据的前提下能够泛化到该任务的假设所占的体积分数实现,其数值随模型需泛化的空间本征维度呈指数增长,但随每个维度的分辨率仅呈多项式增长,这表明需跨多维度泛化的任务比涉及较少维度但细节更丰富的任务困难得多。该方法可应用于计算并比较监督学习、强化学习与元学习任务间的泛化难度。实证表明,该度量能形式化量化直观预期趋势,例如就所需归纳偏置而言,MNIST < CIFAR10 < ImageNet,且完全可观测马尔可夫决策过程(MDP)< 部分可观测MDP。此外,我们还发现复杂图像分类 < 基于简单图像的少样本元学习。该度量提供了定量指标以指导构建需要更大归纳偏置的复杂任务,从而促进开发具有更强泛化能力的更先进架构与学习算法。