NTH

Learning Length-Extrapolatable Recurrent Models

AuthorsHanwen Jiang

September 10, 2026 2 min read
Watch on YouTube
The one-line take

The paper helps recurrent networks learn reliably far beyond the sequence lengths they saw during training by stabilizing how future errors credit earlier states.

Key results

6.58%
Meta-FSA long-context gain

Average Event-CST improvement over BPTT across 16 evaluations from 16K to 128K.

41.66%
MQAR accuracy at 128K

Event-CST accuracy, compared with 25.56% for BPTT.

6.91
EMA replay boundary hooks

Average exposed boundaries per step, reduced from 31 for dense Event-CST.

84.6%
EMA replay throughput

Projected throughput relative to BPTT.

0.00724
Language-model NLL gain

Average full-token NLL improvement across all 14 dataset–length evaluations.

3.518
Books3 CST NLL at 128K

CST result versus 4.415 for the similarly sized Transformer baseline.

What the paper found

This paper argues that recurrent models fail to extrapolate not simply because gradients vanish, but because future losses may deliver unusable credit to earlier hidden states. It introduces Credit Stabilization through Time, or CST, which rescales the norm of state-credit signals during backpropagation without changing their direction, the forward recurrence, or the training objective. Event-CST targets accumulated contraction on controlled tasks, while a symmetric, head-wise controller handles the irregular expansion and contraction found in real text. Using a three-layer Gated DeltaNet trained at 1K tokens, Event-CST improved Meta-FSA accuracy over BPTT by 6.58 percentage points on average across all 16 evaluations from 16K to 128K, and on MQAR it raised 128K accuracy from 25.56% to 41.66%. EMA replay reduced exposed boundary hooks from 31 to 6.91 per step and recovered 84.6% of BPTT throughput. For language modeling, a 123.30M-parameter recurrent model trained on 4K-token sequences and 7.86B training tokens improved full-token negative log-likelihood in all 14 dataset–length evaluations, with an average gain of 0.00724. On Books3 at 128K, CST reached 3.518 NLL, compared with 4.415 for a similarly sized Transformer using dynamic YaRN; the study also uses Qwen2-72B for key-token annotation and discusses how the result complements Mamba-style recurrent architectures. The paper reports an OpenAI Codex-assisted research workflow, but emphasizes that CST is a backward-pass intervention rather than a new forward architecture.

Original abstract

Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
03Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis