Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the major linear-attention families are term-for-term identical with named models from a century of animal learning theory: linear attention implements Hebbian contiguity, DeltaNet implements Rescorla--Wagner error correction, and decay variants such as RetNet implement contiguity with a stimulus trace. This dictionary turns conditioning phenomena into testable statements about the in-context behavior of linear transformers, while distinguishing algebraic consequences from empirical measurements. Algebraically, it yields an exact closed form for Kamin blocking, verified in simulation to $<10^{-7}$ across five learning rates. Empirically, it predicts a dissociation that survives training on generic in-context association: error-correcting attention exhibits cue competition, whereas contiguity-based attention does not. A single state also has two capacity regimes, with measured scaling exponents of 1.22 for faithful retrieval and 1.89 for identification, consistent with linear and near-quadratic predictions. Across the full head grid, retrieval error is governed primarily by total state size rather than its partition across heads, indicating that heads provide capacity rather than redundant copies. We also prove no spontaneous recovery for the analyzed single-state recurrences under cue-orthogonal retention trials; with a never-presented-cue control and probes within the trained positional range, we likewise find no recovery in trained models. Finally, we introduce PH-attention, a Pearce--Hall-inspired rule with an explicit feature-indexed associability state that yields cue-dependent learning rates and is absent from the token-computed gates we compare.
翻译:暂无翻译