CoWindow Attention: Full Causal Coverage Is a Collective Property
AuthorsJingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
Key results
Associative-recall accuracy with full collective coverage.
Dense-attention reference in the matched ablation.
CoWA speedup over FullAttn at 128K tokens.
CoWA speedup over FullAttn at 128K tokens.
Inference decoding speedup over FullAttn at 128K tokens.
Reduction achieved by the 14B CoWA model versus FullAttn.
What the paper found
CoWindow Attention, or CoWA, argues that full causal access does not need to be duplicated in every attention head. Instead, under grouped-query attention, all KV heads share recent-token and prefix-sink windows, while complementary long-range windows partition the remaining history; their union provides full causal coverage without a learned router or indexer. The same position-defined pattern runs during training, prefill, decoding, and tensor-parallel execution on NVIDIA H100 GPUs. In a matched 8K associative-recall ablation, CoWA reaches 89.73% accuracy versus 89.97% for FullAttn, while duplicated long-range windows perform far worse; at 8K, CoWA also substantially outperforms local sliding-window attention. At 128K tokens, CoWA reduces training forward latency by 7.4 times and backward latency by 8.6 times, while decoding is 3.0 times faster than FullAttn; decoding operator memory is 7.6 times lower. Scaling experiments from 0.6B to 14B parameters show nearly identical perplexity to FullAttn, with 28.5% lower total training FLOPs during 32K-context training. The resulting 14B models retain comparable knowledge, reasoning, and RULER retrieval scores, including 66.60 versus 65.84 for FullAttn at 128K after YaRN extrapolation. Compared with dynamic approaches such as DeepSeek’s DSA and MoBA, CoWA gains efficiency from regular position-based sparsity rather than routing overhead, while using the Qwen3 tokenizer in its scaling-law training.
Original abstract
FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.
Block Sparse Attention with Log-Linear Complexity
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu
PISA makes long-context attention more scalable by hierarchically narrowing relevant key blocks instead of comparing every query with every block.