Explore how models select and combine information through attention. Follow research on long contexts, efficient computation, and architectural alternatives.
18 papers · Latest edition October 9, 2026
Where to start
Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
Memory Attention replaces some costly attention computation with reusable token memories, potentially making language models cheaper to run and easier to fit on limited GPU hardware.
The study finds that real paragraph structure is revealed not merely by where attention concentrates, but by how deeply that compression develops across model layers.
This work explains how different ways of parameterizing attention can determine whether large models learn useful structure or remain stuck, using a detailed high-dimensional theory.
This work argues that a power-law, graph-based form of attention can mathematically subsume standard attention and may effectively collapse to it during inference.
This paper makes spiking transformers more spatially aware by combining localized attention with a lightweight pathway that preserves communication across region boundaries.
Tommaso Cerruti, Tim Rieder, George Rowlands, Lingfeng Jin, Imanol Schlag
This paper compares several modern linear-attention designs under one framework and finds that a simple cross-layer value-routing trick can modestly improve performance.
This paper makes linear attention smarter by dynamically deciding how to merge memory states, aiming to keep long-context models efficient without losing important information.
Zhentao Tan, Wei Chen, Jingyi Shen, Yao Liu, Xu Shen, Yue Wu, Jieping Ye
HydraHead rethinks hybrid attention by assigning full attention only to the most retrieval-critical heads, improving long-context performance while keeping compute efficient.
This paper proposes LOCOS, a new way to find attention heads that retrieve meaning rather than just copied words, and shows those heads are crucial for long-context question answering.
This paper introduces a biologically inspired attention mechanism that turns attention scores into stochastic variables to produce uncertainty-aware predictions across several continuous-time tasks.
This paper replaces softmax attention with a hysteresis-based memory model that tracks extrema, aiming to give transformers a more efficient way to remember long histories.