Forget Attention: Importance-Aware Attention Is All You Need
AuthorsSoohyeong Shin, Yeongwook Yang
Resources
SISA is a new attention design that lets state-space models guide which tokens matter inside softmax attention, aiming to combine global retrieval with better prioritization in language models.
Key results
SISA at 152M on 5B SlimPajama tokens reaches 17.3% LAMBADA-greedy, compared with the Transformer’s 13.9% and Mamba-3’s 15.5%.
What the paper found
Forget Attention: Importance-Aware Attention Is All You Need, from Suhyeong Shin and Yeongwook Yang at Kangwon National University, proposes SISA, a score-level fusion method that injects a state space model signal directly into softmax attention instead of keeping attention and SSMs in separate layers or heads like Jamba and Hymba. The key idea is to augment query and key vectors with a low-dimensional SSM-derived importance channel, so the attention logit becomes content similarity plus a learned inner-product bias encoding cumulative decay and data-dependent rotation, then compute everything with one standard PyTorch SDPA call compatible with FlashAttention. On 5B tokens of SlimPajama-6B, SISA at 152M parameters reaches 17.3% on LAMBADA-greedy versus 13.9% for the Transformer and 15.5% for Mamba-3, while holding NIAH at 100% from step 1K, about 7× faster than the Transformer’s retrieval convergence; at 50M it also beats the Transformer when ds is tuned to 64, and at 369M Mamba-3 leads LAMBADA but SISA preserves perfect NIAH and stock-kernel execution. The paper’s main novelty is architectural: it defines a third hybrid design axis, score-level fusion, and shows that the optimal SSM channel size ds is non-monotonic across scale, reflecting a trade-off between retrieval bias and FFN capacity.
Original abstract
Combining attention's global retrieval with the sequential importance signal of state space models (SSMs) is the open challenge of hybrid language modeling. Transformers see everywhere but cannot prioritize; SSMs know what matters but cannot revisit. Existing hybrids -- Jamba (block level) and Hymba (head level) -- place the two in separate compartments, so neither informs the other during the attention computation itself. We propose SISA (SSM-Informed Softmax Attention), which adds an SSM-derived importance term directly inside the attention score and realizes the full operation as a single SDPA call on augmented query/key vectors -- no recurrent state, no custom kernel. At 152M / 5B tokens, SISA reaches LAMBADA-greedy 17.3% (vs. Transformer 13.9 and Mamba-3 15.5) and attains NIAH 100% from step 1K, 7x faster than Transformer's retrieval convergence; at 369M, Mamba-3 leads LAMBADA while SISA preserves perfect NIAH and stock-SDPA execution. SISA thus defines a third design axis for SSM-attention hybrids -- score-level fusion -- beyond the block-level and head-level paradigms that have dominated the field.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.