NTH

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

AuthorsXiang Hu, Xinyu Wei, Hao Gu, Minshen Zhang, Tian Liang, Huayang Li, Lei Zhu, Yan Wang, Sirui Han, Yushi Bai, Kewei Tu, Haitao Mi, Leo Liang

July 8, 2026 2 min read
Watch on YouTube
The one-line take

HiLS Attention is a new way to make LLMs handle much longer contexts by learning which chunks to attend to, cutting compute while keeping performance strong.

Key results

8K
Training context

Base continued-training length for the 345M and 7B experiments

4M
Extrapolation length

HiLS-Attn-HoPE RULER needle retrieval evaluation length

90%
Retrieval accuracy

HiLS-Attn-HoPE maintains over 90% on needle retrieval at ultra-long context

13.5x
Prefill speedup

HiLS-Attention vs full attention at 512K context on NVIDIA H800

15.7x
Decode speedup

HiLS-Attention vs full attention at 512K context on NVIDIA H800

50B
Continual training tokens

Tokens used to convert the 7B Olmo3-1025 checkpoint to HiLS-Attn

What the paper found

Hierarchical Sparse Attention Done Right Toward Infinite Context Modeling, from Tencent Hunyuan with collaborators at ShanghaiTech University, HKUST, and UC San Diego, argues that chunk-wise sparse attention fails mainly because chunk selection is learned too weakly. The paper introduces HiLS-Attention, which factorizes attention into inter-chunk routing and intra-chunk token attention, then makes routing end-to-end trainable under the language-modeling loss by using a landmark token per chunk, a Taylor-linearized chunk-mass surrogate, and a low-rank query calibration module. On 345M models trained at 8K context, HiLS matches full attention perplexity at 8K and extrapolates to 4M context with RULER needle retrieval still above 90%, while outperforming prior sparse methods such as NSA, DashAttention, and InfLLM v2 on exact-match retrieval. Inference is also much cheaper: on a single NVIDIA H800, HiLS reaches parity around 16K context and at 512K is 13.5× faster in prefill and 15.7× faster per-token decode than full attention. At larger scale, the authors convert the 7B Olmo3-1025 checkpoint using only 50B continual-training tokens and show stronger LongBench-v1 results than the YaRN-extended baseline, with an overall score of 33.2 versus 31.7. The core claim is that accurately learned sparse retrieval can preserve short-context quality while enabling ultra-long, potentially infinite, context modeling without quadratic attention cost.

Original abstract

Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of dense attention. Chunk-wise sparse attention offers a promising alternative, but all existing methods fall short of full attention because of their inaccurate chunk selection. We propose Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling (LM) loss. HiLS factorizes attention hierarchically: each query performs attention independently with each retrieved chunk to extract chunk-specific information, and the resulting outputs are fused according to chunk retrieval scores. By incorporating retrieval scores into the forward attention computation, HiLS optimizes them directly with the LM loss, enabling end-to-end retrieval learning and native sparse training. Experimental results show that HiLS-Attention achieves performance comparable to, and in some cases better than, full attention at in-domain context lengths. Meanwhile, HiLS-Attention extrapolates more than $64\times$ the training context length with 90% retrieval accuracy, far beyond full attention. Moreover, existing full-attention models can be converted to HiLS-Attention with lightweight continued pretraining, preserving in-domain performance while acquiring ultra-long-context extrapolation. Together with its sparse KV access and computation, HiLS-Attention breaks the usual efficiency-performance trade-off, enabling long-context LLMs that are both more efficient and more effective on general long-context tasks than their full-attention counterparts.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis