NTH

Large-Step Training Dynamics of a Two-Factor Linear Transformer Model

AuthorsKrishnakumar Balasubramanian

May 22, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that, in a simplified transformer model, large learning rates can make training settle into cycles or chaos instead of cleanly learning, revealing new stability thresholds for transformer dynamics.

Key results

2 √2 − 2 ≈ 0.828
Monotone error threshold

Exact boundary on the balanced cubic dynamics between uniform one-step error contraction and possible catapult growth.

1
Stability flip threshold

At this effective step size, the zero-error fixed point loses local stability and the first period-two orbit appears.

2
Divergence threshold

Above this effective step size, the balanced scalar dynamics diverge and the two-dimensional system has escaping branches.

0 < µ < 2
Invariant Chebyshev ellipse

For this parameter range, the full two-dimensional map has an explicit invariant ellipse carrying angle-tripling dynamics.

1 < µ < 2
Transverse attraction window

Balanced invariant measures are transversely attracting only in this range, even though the off-balanced ellipse remains repelling.

µ(ξ 2 + 2ξ) > 2
Mini-batch crossing condition

A single mini-batch with parameter ν = µ + ξ can push the iterate from the full-batch interior to the exterior of the separatrix.

What the paper found

This paper derives an exact finite-step training model for a one-prompt linear self-attention transformer and shows that large learning rates can create new attractors rather than simply speeding convergence. After rescaling, gradient descent on the reduced one-prompt loss becomes the two-factor map Φµ(a,b) = (a − (ab − µ)b, b − (ab − µ)a), equivalent to unit-step gradient descent on ℓµ(a,b)=1/2(ab−µ)^2 with effective step size µ = 2η|yq|√κ. On the balanced slice a=b, the dynamics reduce to the cubic family studied in quadratic regression, giving sharp thresholds at µ = 2√2 − 2 ≈ 0.828 for monotone error contraction, µ = 1 for loss of stability and a flip bifurcation, and µ = 2 for divergence. The novel result is the full two-dimensional analysis: for 0 < µ < 2, the map has an explicit invariant Chebyshev ellipse Eµ, on which the error evolves as e↦e^3−3e, i.e. angle tripling e = 2 cos θ ⇒ e+ = 2 cos 3θ. This ellipse is not an attractor but a transversely repelling separatrix, while balanced invariant measures are transversely attracting only for 1 < µ < 2. The paper also proves that the strict interior contains no hidden off-balanced recurrent dynamics beyond algebraic landing sets. For mini-batch gradient descent, each batch induces its own Φν, so training is random switching between distinct separatrices; a single atypical batch can push the iterate across the full-batch boundary. The main implication is mechanistic: in linear-transformer training, increasing η can move optimization into catapult, periodic, chaotic, or divergent regimes, and minibatch noise can destabilize a population-stable solution by exposing modes with batch-specific effective µ past the thresholds.

Original abstract

Gradient-flow analyses show that simplified linear transformers can learn the in-context linear-regression algorithm, but they do not explain the finite-step behavior of gradient descent at large learning rates. Motivated by empirical work on high-learning-rate transformer instabilities and by the cubic-map phase diagram for quadratic regression, we study an exactly reducible one-prompt linear-transformer training problem. After normalization, the dynamics reduce to a two-factor product map with an effective step-size parameter \(μ\). On the balanced slice, this map recovers the known scalar cubic transition from monotone convergence to catapult convergence, periodic and chaotic bounded nonconvergence, and divergence. We then analyze the full two-dimensional system and show that, for \(0<μ<2\), it has an explicit invariant Chebyshev ellipse separating forward-invariant regions; this ellipse carries off-balanced chaotic dynamics but is transversely repelling, while balanced scalar attractors can be transversely attracting. These results show that large constant learning rates can change the training attractor of the learned transformer rather than merely accelerating convergence: beyond sharp stability thresholds, finite-step training may settle into cycles, bounded chaos, or divergence instead of a single in-context linear-regression solution. We also discuss the consequences for mini-batch gradient descent based training methods.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis