Kalman Delta Networks: Uncertainty-aware Associative Memory
AuthorsNgoc Bui, Tinglin Huang, Rex Ying
Resources
This work makes linear attention more adaptive by letting its memory track not only what it knows, but also how confident it is.
Key results
Parameter count for the smaller controlled pretraining setup.
FineWeb-Edu tokens used at the 750M scale.
Mean six-task zero-shot accuracy, compared with 53.87 for KDA.
Parameter count for the larger controlled pretraining setup.
Mean six-task zero-shot accuracy after training on 100B tokens.
What the paper found
Kalman Delta Networks, or KDNs, recast recurrent associative memory as a linear–Gaussian state-space model. Instead of choosing a write gate solely from the current token, the model propagates both a key–value memory estimate and its uncertainty, then uses a Kalman gain to weight each residual update according to accumulated evidence and observation noise. Because exact covariance tracking requires a dense Riccati recursion that is unsuitable for parallel GPU scans, the paper introduces Diagonal KDN, which tracks one uncertainty value per key channel through online mean-field variational inference, and Isotropic KDN, which keeps one uncertainty scalar per head. Both uncertainty recurrences are Möbius maps, preserving associative scan computation; their auxiliary uncertainty state is O(dk) and O(1), respectively. In controlled FineWeb-Edu pretraining, KDN is compared with DeltaNet, KDA, GDN-2, and Mamba-3. At 750M parameters trained on 50B tokens, Diagonal KDN reaches 54.97 mean accuracy across six zero-shot tasks, versus 53.87 for KDA. At 1.3B parameters trained on 100B tokens, it reaches 60.45, while also achieving WikiText and LAMBADA perplexities of 15.04 and 9.75. Diagonal KDN also leads recurrent baselines on RULER retrieval, indicating that uncertainty-aware writes better protect associations from interference during long-context recall.
Original abstract
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.