Despite their omnipresence in modern NLP, characterizing the computational power of transformer neural nets remains an interesting open question. We prove that transformers whose arithmetic precision is logarithmic in the number of input tokens (and whose feedforward nets are computable using space linear in their input) can be simulated by constant-depth logspace-uniform threshold circuits. This provides insight on the power of transformers using known results in complexity theory. For example, if $\mathsf L \neq \mathsf P$ (i.e., not all poly-time problems can be solved using logarithmic space), then transformers cannot even accurately solve linear equalities or check membership in an arbitrary context-free grammar with empty productions. Our result intuitively emerges from the transformer architecture's high parallelizability. We thus speculatively introduce the idea of a fundamental parallelism tradeoff: any model architecture as parallelizable as the transformer will obey limitations similar to it. Since parallelism is key to training models at massive scale, this suggests a potential inherent weakness of the scaling paradigm.
翻译:尽管Transformer在现代NLP中无处不在,但其神经网络的的计算能力表征仍是一个有趣的开放性问题。我们证明,算术精度为输入token数的对数(且其前馈网络的计算空间与输入长度成线性关系)的Transformer,可由常数深度对数空间一致阈值电路模拟。这一结果借助复杂性理论的已知结论,揭示了Transformer的能力边界。例如,若$\mathsf L \neq \mathsf P$(即并非所有多项式时间问题都能在对数空间中求解),则Transformer甚至无法准确求解线性等式,或判定任意含空产生式的上下文无关文法的成员关系。我们的结论直观源于Transformer架构的高度可并行性。由此我们推测性地提出基本并行性权衡的概念:任何具有类似Transformer并行性的模型架构,都将服从类似的限制。由于并行性是大规模训练模型的关键,这暗示了规模化范式潜在的根本性缺陷。