Learning Length-Extrapolatable Recurrent Models
AuthorsHanwen Jiang
Resources
The paper helps recurrent networks learn reliably far beyond the sequence lengths they saw during training by stabilizing how future errors credit earlier states.
Key results
Average Event-CST improvement over BPTT across 16 evaluations from 16K to 128K.
Event-CST accuracy, compared with 25.56% for BPTT.
Average exposed boundaries per step, reduced from 31 for dense Event-CST.
Projected throughput relative to BPTT.
Average full-token NLL improvement across all 14 dataset–length evaluations.
CST result versus 4.415 for the similarly sized Transformer baseline.
What the paper found
This paper argues that recurrent models fail to extrapolate not simply because gradients vanish, but because future losses may deliver unusable credit to earlier hidden states. It introduces Credit Stabilization through Time, or CST, which rescales the norm of state-credit signals during backpropagation without changing their direction, the forward recurrence, or the training objective. Event-CST targets accumulated contraction on controlled tasks, while a symmetric, head-wise controller handles the irregular expansion and contraction found in real text. Using a three-layer Gated DeltaNet trained at 1K tokens, Event-CST improved Meta-FSA accuracy over BPTT by 6.58 percentage points on average across all 16 evaluations from 16K to 128K, and on MQAR it raised 128K accuracy from 25.56% to 41.66%. EMA replay reduced exposed boundary hooks from 31 to 6.91 per step and recovered 84.6% of BPTT throughput. For language modeling, a 123.30M-parameter recurrent model trained on 4K-token sequences and 7.86B training tokens improved full-token negative log-likelihood in all 14 dataset–length evaluations, with an average gain of 0.00724. On Books3 at 128K, CST reached 3.518 NLL, compared with 4.415 for a similarly sized Transformer using dynamic YaRN; the study also uses Qwen2-72B for key-token annotation and discusses how the result complements Mamba-style recurrent architectures. The paper reports an OpenAI Codex-assisted research workflow, but emphasizes that CST is a backward-pass intervention rather than a new forward architecture.
Original abstract
Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.