Large-Step Training Dynamics of a Two-Factor Linear Transformer Model
AuthorsKrishnakumar Balasubramanian
Resources
This paper shows that, in a simplified transformer model, large learning rates can make training settle into cycles or chaos instead of cleanly learning, revealing new stability thresholds for transformer dynamics.
Key results
Exact boundary on the balanced cubic dynamics between uniform one-step error contraction and possible catapult growth.
At this effective step size, the zero-error fixed point loses local stability and the first period-two orbit appears.
Above this effective step size, the balanced scalar dynamics diverge and the two-dimensional system has escaping branches.
For this parameter range, the full two-dimensional map has an explicit invariant ellipse carrying angle-tripling dynamics.
Balanced invariant measures are transversely attracting only in this range, even though the off-balanced ellipse remains repelling.
A single mini-batch with parameter ν = µ + ξ can push the iterate from the full-batch interior to the exterior of the separatrix.
What the paper found
This paper derives an exact finite-step training model for a one-prompt linear self-attention transformer and shows that large learning rates can create new attractors rather than simply speeding convergence. After rescaling, gradient descent on the reduced one-prompt loss becomes the two-factor map Φµ(a,b) = (a − (ab − µ)b, b − (ab − µ)a), equivalent to unit-step gradient descent on ℓµ(a,b)=1/2(ab−µ)^2 with effective step size µ = 2η|yq|√κ. On the balanced slice a=b, the dynamics reduce to the cubic family studied in quadratic regression, giving sharp thresholds at µ = 2√2 − 2 ≈ 0.828 for monotone error contraction, µ = 1 for loss of stability and a flip bifurcation, and µ = 2 for divergence. The novel result is the full two-dimensional analysis: for 0 < µ < 2, the map has an explicit invariant Chebyshev ellipse Eµ, on which the error evolves as e↦e^3−3e, i.e. angle tripling e = 2 cos θ ⇒ e+ = 2 cos 3θ. This ellipse is not an attractor but a transversely repelling separatrix, while balanced invariant measures are transversely attracting only for 1 < µ < 2. The paper also proves that the strict interior contains no hidden off-balanced recurrent dynamics beyond algebraic landing sets. For mini-batch gradient descent, each batch induces its own Φν, so training is random switching between distinct separatrices; a single atypical batch can push the iterate across the full-batch boundary. The main implication is mechanistic: in linear-transformer training, increasing η can move optimization into catapult, periodic, chaotic, or divergent regimes, and minibatch noise can destabilize a population-stable solution by exposing modes with batch-specific effective µ past the thresholds.
Original abstract
Gradient-flow analyses show that simplified linear transformers can learn the in-context linear-regression algorithm, but they do not explain the finite-step behavior of gradient descent at large learning rates. Motivated by empirical work on high-learning-rate transformer instabilities and by the cubic-map phase diagram for quadratic regression, we study an exactly reducible one-prompt linear-transformer training problem. After normalization, the dynamics reduce to a two-factor product map with an effective step-size parameter \(μ\). On the balanced slice, this map recovers the known scalar cubic transition from monotone convergence to catapult convergence, periodic and chaotic bounded nonconvergence, and divergence. We then analyze the full two-dimensional system and show that, for \(0<μ<2\), it has an explicit invariant Chebyshev ellipse separating forward-invariant regions; this ellipse carries off-balanced chaotic dynamics but is transversely repelling, while balanced scalar attractors can be transversely attracting. These results show that large constant learning rates can change the training attractor of the learned transformer rather than merely accelerating convergence: beyond sharp stability thresholds, finite-step training may settle into cycles, bounded chaos, or divergence instead of a single in-context linear-regression solution. We also discuss the consequences for mini-batch gradient descent based training methods.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.