To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention. In the zero-temperature limit, we show that fixed-key top-$1$ routing is exactly represented by a power diagram in query space, while an auxiliary log-lifted value parameterization yields a vector-valued tropical rational representation. For Multi-Head Self-Attention (MHSA) with sequence length $N$ and $H$ attention heads, the joint routing geometry is encoded by Minkowski sums of headwise Newton polytopes, giving an $\mathcal{O}(N^H)$ universal bound that sharpens to $\mathcal{O}((HN)^{d_{\mathrm{model}}-1})$ once the number of heads reaches the intrinsic dimension $d_{\mathrm{model}}$. Extending this analysis across depth $L$, we derive the first tight asymptotic bounds on the number of linear regions in transformers ($Θ\!\left(N^{\min\{H,d_{\mathrm{model}}-1\}L}\right)$). We further show that finite-temperature softmax preserves the top-$1$ routing structure and admits exponentially decaying local approximation and differential bounds away from routing boundaries.
翻译:暂无翻译