NTH

Rethinking Attention Locality in Spiking Transformers

AuthorsZeqi Zheng, Zizheng Zhu, Yuping Yan, Wenxuan Pan, Zhaofei Yu, Yaochu Jin

August 15, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes spiking transformers more spatially aware by combining localized attention with a lightweight pathway that preserves communication across region boundaries.

Key results

7
Evaluation datasets

Static and neuromorphic datasets spanning classification, detection, and segmentation.

78.62%
Spikingformer ImageNet-1K top-1

SCLA-BCP result, compared with the 77.64% baseline.

9.50%
COCO 2017 mAP@50 improvement

Maximum improvement over the baseline on Microsoft COCO 2017.

3.42%
ADE20K mIoU improvement

Maximum improvement over the baseline on semantic segmentation.

0.03M
Spikingformer parameter overhead

Additional parameters for the ImageNet-1K Spikingformer configuration.

What the paper found

Rethinking Attention Locality in Spiking Transformers argues that making attention computationally cheaper does not guarantee spatially local token interactions. In Softmax-free Spiking Self-Attention, offset-based grouping in LSSA can connect distant pixels, while uniformly applying LRF-SSA produces inconsistent Mean Attention Distance across layers and architectures. The paper introduces Spatially Contiguous Local Attention with Boundary Continuity Pathway, or SCLA-BCP: SCLA partitions feature maps into non-overlapping, spatially contiguous regions and performs attention within each region, while BCP uses a lightweight 3×3 depth-wise convolution, batch normalization, and max-pooling pathway to exchange information across region boundaries. A hierarchical deployment strategy applies SCLA-BCP to the first half of SSA layers in ViT-like models and to Stages 1–2 in hierarchical models. Across 7 static and neuromorphic datasets, including ImageNet-1K, CIFAR10-DVS, N-Caltech101, Microsoft COCO 2017, and ADE20K, the method consistently improves classification, detection, and segmentation. On ImageNet-1K, Spikingformer reaches 78.62% top-1 accuracy versus 77.64% for its baseline; on Microsoft COCO 2017, it improves mAP@50 by up to 9.50%; and on ADE20K, it improves mIoU by up to 3.42%. The gains require limited overhead, including just 0.03M additional parameters in the Spikingformer ImageNet-1K configuration, while MAD measurements and ablations confirm that contiguous regional attention, combined with boundary compensation, produces genuinely more localized interactions.

Original abstract

Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computation, yet their Softmax-free Spiking Self-Attention (SSA) struggles to establish spatially localized token interactions. Although existing locality-enhanced SSA methods improve accuracy, it remains unclear whether they consistently induce spatial locality across layers and different Spiking Transformer architectures. Through Mean Attention Distance (MAD) analysis, we reveal that computational locality does not necessarily translate into spatial locality and show that uniformly applying the same locality enhancement overlooks architecture-dependent deployment requirements. Motivated by these observations, we propose Spatially Contiguous Local Attention with Boundary Continuity Pathway (SCLA-BCP). SCLA computes attention within non-overlapping regions of spatially adjacent tokens, while BCP facilitates cross-boundary information exchange through a lightweight convolutional pathway. Furthermore, we develop a hierarchical locality deployment strategy to effectively apply SCLA-BCP across the two major Spiking Transformer architectures. Extensive experiments on seven static and neuromorphic datasets covering classification, detection, and segmentation demonstrate consistent improvements with limited parameter and energy overhead. Notably, our approach improves mAP@50 by up to 9.50% on COCO 2017 and mIoU by up to 3.42% on ADE20K. Visualizations, MAD analysis, and ablation studies further validate its effectiveness.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis