NTH

CoWindow Attention: Full Causal Coverage Is a Collective Property

AuthorsJingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

October 8, 2026 3 min read
Watch on YouTube
The one-line take

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Key results

89.73
CoWA 8K recall accuracy

Associative-recall accuracy with full collective coverage.

89.97
FullAttn 8K recall accuracy

Dense-attention reference in the matched ablation.

7.4
Training forward latency speedup

CoWA speedup over FullAttn at 128K tokens.

8.6
Training backward latency speedup

CoWA speedup over FullAttn at 128K tokens.

3.0
Decoding latency speedup

Inference decoding speedup over FullAttn at 128K tokens.

28.5%
32K training FLOP reduction

Reduction achieved by the 14B CoWA model versus FullAttn.

What the paper found

CoWindow Attention, or CoWA, argues that full causal access does not need to be duplicated in every attention head. Instead, under grouped-query attention, all KV heads share recent-token and prefix-sink windows, while complementary long-range windows partition the remaining history; their union provides full causal coverage without a learned router or indexer. The same position-defined pattern runs during training, prefill, decoding, and tensor-parallel execution on NVIDIA H100 GPUs. In a matched 8K associative-recall ablation, CoWA reaches 89.73% accuracy versus 89.97% for FullAttn, while duplicated long-range windows perform far worse; at 8K, CoWA also substantially outperforms local sliding-window attention. At 128K tokens, CoWA reduces training forward latency by 7.4 times and backward latency by 8.6 times, while decoding is 3.0 times faster than FullAttn; decoding operator memory is 7.6 times lower. Scaling experiments from 0.6B to 14B parameters show nearly identical perplexity to FullAttn, with 28.5% lower total training FLOPs during 32K-context training. The resulting 14B models retain comparable knowledge, reasoning, and RULER retrieval scores, including 66.60 versus 65.84 for FullAttn at 128K after YaRN extrapolation. Compared with dynamic approaches such as DeepSeek’s DSA and MoBA, CoWA gains efficiency from regular position-based sparsity rather than routing overhead, while using the Qwen3 tokenizer in its scaling-law training.

Original abstract

FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
03Attention

Block Sparse Attention with Log-Linear Complexity

Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu

PISA makes long-context attention more scalable by hierarchically narrowing relevant key blocks instead of comparing every query with every block.

Read analysis