NTH

Kalman Delta Networks: Uncertainty-aware Associative Memory

AuthorsNgoc Bui, Tinglin Huang, Rex Ying

September 9, 2026 3 min read
Watch on YouTube
The one-line take

This work makes linear attention more adaptive by letting its memory track not only what it knows, but also how confident it is.

Key results

750M
750M training scale

Parameter count for the smaller controlled pretraining setup.

50B
750M training tokens

FineWeb-Edu tokens used at the 750M scale.

54.97
750M Diagonal KDN mean accuracy

Mean six-task zero-shot accuracy, compared with 53.87 for KDA.

1.3B
1.3B training scale

Parameter count for the larger controlled pretraining setup.

60.45
1.3B Diagonal KDN mean accuracy

Mean six-task zero-shot accuracy after training on 100B tokens.

What the paper found

Kalman Delta Networks, or KDNs, recast recurrent associative memory as a linear–Gaussian state-space model. Instead of choosing a write gate solely from the current token, the model propagates both a key–value memory estimate and its uncertainty, then uses a Kalman gain to weight each residual update according to accumulated evidence and observation noise. Because exact covariance tracking requires a dense Riccati recursion that is unsuitable for parallel GPU scans, the paper introduces Diagonal KDN, which tracks one uncertainty value per key channel through online mean-field variational inference, and Isotropic KDN, which keeps one uncertainty scalar per head. Both uncertainty recurrences are Möbius maps, preserving associative scan computation; their auxiliary uncertainty state is O(dk) and O(1), respectively. In controlled FineWeb-Edu pretraining, KDN is compared with DeltaNet, KDA, GDN-2, and Mamba-3. At 750M parameters trained on 50B tokens, Diagonal KDN reaches 54.97 mean accuracy across six zero-shot tasks, versus 53.87 for KDA. At 1.3B parameters trained on 100B tokens, it reaches 60.45, while also achieving WikiText and LAMBADA perplexities of 15.04 and 9.75. Diagonal KDN also leads recurrent baselines on RULER retrieval, indicating that uncertainty-aware writes better protect associations from interference during long-context recall.

Original abstract

Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis