Fast Weight Attention for Continual Learning
AuthorsYifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
Resources
This paper turns attention into an online learning memory, showing how fast-weight updates can help sequence models retain and use information over long contexts.
Key results
Representative Falcon and baseline models used for the language-model comparisons.
Token budget used to train the 124M–130M-parameter language models.
Validation perplexity achieved by the strongest evaluated regression variant.
Average accuracy across eight downstream tasks.
Mean teacher-forced accuracy on variable-digit addition across 33–48 digits.
Comparison score on the same variable-digit addition extrapolation benchmark.
What the paper found
Fast Weight Attention for Continual Learning treats a recurrent memory matrix as an online learner rather than a passive cache. Its central correction is temporal alignment: under read-after-write autoregression, each newly observed value v_t is written using the prefix feature ϕ(k_{t−1}), not the commonly used same-step feature ϕ(k_t). From this causal pairing, the paper derives Falcon-1, Falcon-2, and Falcon-3 regression updates using normalized least-mean-squares steps, explicit ridge-based forgetting, per-channel plasticity, and sliding-window rehearsal; Falcon-1A, Falcon-2A, and Falcon-3A provide corresponding inner-product writes. WY, triangular-solve, and ParallelFlow formulations enable chunk-parallel training while retaining fixed-size recurrent inference, complementing architectures such as Mamba-2, DeltaNet, Gated DeltaNet, and Transformer models. On FineWeb-Edu, 124M–130M-parameter models trained with a 50B-token budget achieved perplexity 17.10 for Falcon-1.3, compared with 17.32 for Gated DeltaNet, while Falcon-1A.2 reached a 49.30 zero-shot average and Falcon-1.3 reached 49.54 one-shot accuracy. On variable-digit addition trained through width 32 and evaluated across 33–48 digits, Falcon-3A.3 achieved 87.2 mean accuracy, exceeding Falcon-1A.3 at 85.9 and demonstrating improved length extrapolation. The implementation uses RMSNorm-based feature scaling and was evaluated on NVIDIA H100 or H200 GPUs; the work includes a ByteDance Seed contribution.
Original abstract
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step $t$ is the prefix-aligned pair $(\mathbf{x}_t,\mathbf{y}_t)=(φ(\mathbf{k}_{t-1}),\mathbf{v}_t)$. The common same-step association $(φ(\mathbf{k}_t),\mathbf{v}_t)$ remains causal, but optimizes a different internal objective. We derive normalized first-order updates for squared-error regression and negative inner-product objectives. The regression family comprises Falcon-1 (a scalar NLMS update), Falcon-2 (its per-column extension), and Falcon-3 (a sliding-window mini-batch update); Falcon-1A/Falcon-2A/Falcon-3A are the corresponding inner-product variants. We provide recurrent, masked-parallel, and chunk-parallel forms, together with numerically stable positive-decay renormalization. Representative variants remain competitive in language modeling and improve length extrapolation on variable-digit addition. This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.