NTH
Research collection

Attention Mechanisms research

Explore how models select and combine information through attention. Follow research on long contexts, efficient computation, and architectural alternatives.

18 papers · Latest edition October 9, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Attention Mechanisms papers

Newest editions first.

01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis
04Attention

Block Sparse Attention with Log-Linear Complexity

Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu

PISA makes long-context attention more scalable by hierarchically narrowing relevant key blocks instead of comparing every query with every block.

Read analysis
05Attention

Memory Attention

Jiale Kang

Memory Attention replaces some costly attention computation with reusable token memories, potentially making language models cheaper to run and easier to fit on limited GPU hardware.

Read analysis
07Attention

High-Dimensional Learning Dynamics of Attention-Indexed Models

Yizhou Xu, Margarita Sagitova, Lenka Zdeborová, Florent Krzakala

This work explains how different ways of parameterizing attention can determine whether large models learn useful structure or remain stuck, using a detailed high-dimensional theory.

Read analysis
09Attention

Liquid Gated Attention

Yiheng Jiang, Yuanbo Xu, Yongjian Yang

LGA is an efficient attention mechanism that models irregular, long-range time-series dynamics without sequential numerical solvers.

Read analysis
12Attention

Rethinking Attention Locality in Spiking Transformers

Zeqi Zheng, Zizheng Zhu, Yuping Yan, Wenxuan Pan, Zhaofei Yu, Yaochu Jin

This paper makes spiking transformers more spatially aware by combining localized attention with a lightweight pathway that preserves communication across region boundaries.

Read analysis
14Attention

Dynamic Linear Attention

Xin Wang, Hui Shen, Boyuan Zheng, Xueshen Liu, Minkyoung Cho, Zhongwei Wan, Zesen Zhao, Zhuoqing Mao, Shen Yan, Mi Zhang

This paper makes linear attention smarter by dynamically deciding how to merge memory states, aiming to keep long-context models efficient without losing important information.

Read analysis
16Attention

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini

This paper proposes LOCOS, a new way to find attention heads that retrieve meaning rather than just copied words, and shows those heads are crucial for long-context question answering.

Read analysis