A common lens to theoretically study neural net architectures is to analyze the functions they can approximate. However, constructions from approximation theory may be unrealistic and therefore less meaningful. For example, a common unrealistic trick is to encode target function values using infinite precision. To address these issues, this work proposes a formal definition of statistically meaningful (SM) approximation which requires the approximating network to exhibit good statistical learnability. We study SM approximation for two function classes: boolean circuits and Turing machines. We show that overparameterized feedforward neural nets can SM approximate boolean circuits with sample complexity depending only polynomially on the circuit size, not the size of the network. In addition, we show that transformers can SM approximate Turing machines with computation time bounded by $T$ with sample complexity polynomial in the alphabet size, state space size, and $\log (T)$. We also introduce new tools for analyzing generalization which provide much tighter sample complexities than the typical VC-dimension or norm-based bounds, which may be of independent interest.
翻译:一种常见的理论分析神经网络架构的视角是研究它们能近似哪些函数。然而,近似理论中的构造可能不切实际,因此意义有限。例如,一个常见的不切实际技巧是使用无限精度编码目标函数值。为解决这些问题,本文提出统计有意义(SM)近似的形式化定义,要求近似网络展现出良好的统计可学习性。我们研究了两类函数(布尔电路和图灵机)的SM近似。结果表明,过参数化的前馈神经网络能够SM近似布尔电路,其样本复杂度仅与电路规模呈多项式关系,而非网络规模。此外,我们证明Transformer能够SM近似计算时间受$T$限制的图灵机,样本复杂度与字母表大小、状态空间大小及$\log(T)$呈多项式关系。我们还引入了新的泛化分析工具,相比典型的VC维或范数界方法能提供更紧致的样本复杂度界,这些工具本身可能具有独立研究价值。