In order to better understand feature learning in neural networks, we propose a framework for understanding linear models in tangent feature space where the features are allowed to be transformed during training. We consider linear transformations of features, resulting in a joint optimization over parameters and transformations with a bilinear interpolation constraint. We show that this optimization problem has an equivalent linearly constrained optimization with structured regularization that encourages approximately low rank solutions. Specializing to neural network structure, we gain insights into how the features and thus the kernel function change, providing additional nuance to the phenomenon of kernel alignment when the target function is poorly represented using tangent features. In addition to verifying our theoretical observations in real neural networks on a simple regression problem, we empirically show that an adaptive feature implementation of tangent feature classification has an order of magnitude lower sample complexity than the fixed tangent feature model on MNIST and CIFAR-10.
翻译:为了更好地理解神经网络中的特征学习,我们提出一个框架,用于理解切空间特征空间中的线性模型——在该空间中,特征可在训练过程中被变换。我们考虑特征的线性变换,并引入一种带有双线性插值约束的参数与特征联合优化问题。我们证明,该优化问题等价于一个具有结构化正则化的线性约束优化问题,该正则化鼓励得到近似低秩的解。将这一框架特化到神经网络结构,我们得以洞察特征及其核函数的变化规律,从而为目标函数在切空间中表示不佳时核对齐现象提供更细致的解释。除了在简单回归问题中通过真实神经网络验证理论观察外,我们还在MNIST和CIFAR-10数据集上通过实验表明,自适应特征切空间分类器的样本复杂度比固定切空间特征模型低一个数量级。