Rethinking Attention Locality in Spiking Transformers
AuthorsZeqi Zheng, Zizheng Zhu, Yuping Yan, Wenxuan Pan, Zhaofei Yu, Yaochu Jin
Resources
This paper makes spiking transformers more spatially aware by combining localized attention with a lightweight pathway that preserves communication across region boundaries.
Key results
Static and neuromorphic datasets spanning classification, detection, and segmentation.
SCLA-BCP result, compared with the 77.64% baseline.
Maximum improvement over the baseline on Microsoft COCO 2017.
Maximum improvement over the baseline on semantic segmentation.
Additional parameters for the ImageNet-1K Spikingformer configuration.
What the paper found
Rethinking Attention Locality in Spiking Transformers argues that making attention computationally cheaper does not guarantee spatially local token interactions. In Softmax-free Spiking Self-Attention, offset-based grouping in LSSA can connect distant pixels, while uniformly applying LRF-SSA produces inconsistent Mean Attention Distance across layers and architectures. The paper introduces Spatially Contiguous Local Attention with Boundary Continuity Pathway, or SCLA-BCP: SCLA partitions feature maps into non-overlapping, spatially contiguous regions and performs attention within each region, while BCP uses a lightweight 3×3 depth-wise convolution, batch normalization, and max-pooling pathway to exchange information across region boundaries. A hierarchical deployment strategy applies SCLA-BCP to the first half of SSA layers in ViT-like models and to Stages 1–2 in hierarchical models. Across 7 static and neuromorphic datasets, including ImageNet-1K, CIFAR10-DVS, N-Caltech101, Microsoft COCO 2017, and ADE20K, the method consistently improves classification, detection, and segmentation. On ImageNet-1K, Spikingformer reaches 78.62% top-1 accuracy versus 77.64% for its baseline; on Microsoft COCO 2017, it improves mAP@50 by up to 9.50%; and on ADE20K, it improves mIoU by up to 3.42%. The gains require limited overhead, including just 0.03M additional parameters in the Spikingformer ImageNet-1K configuration, while MAD measurements and ablations confirm that contiguous regional attention, combined with boundary compensation, produces genuinely more localized interactions.
Original abstract
Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computation, yet their Softmax-free Spiking Self-Attention (SSA) struggles to establish spatially localized token interactions. Although existing locality-enhanced SSA methods improve accuracy, it remains unclear whether they consistently induce spatial locality across layers and different Spiking Transformer architectures. Through Mean Attention Distance (MAD) analysis, we reveal that computational locality does not necessarily translate into spatial locality and show that uniformly applying the same locality enhancement overlooks architecture-dependent deployment requirements. Motivated by these observations, we propose Spatially Contiguous Local Attention with Boundary Continuity Pathway (SCLA-BCP). SCLA computes attention within non-overlapping regions of spatially adjacent tokens, while BCP facilitates cross-boundary information exchange through a lightweight convolutional pathway. Furthermore, we develop a hierarchical locality deployment strategy to effectively apply SCLA-BCP across the two major Spiking Transformer architectures. Extensive experiments on seven static and neuromorphic datasets covering classification, detection, and segmentation demonstrate consistent improvements with limited parameter and energy overhead. Notably, our approach improves mAP@50 by up to 9.50% on COCO 2017 and mIoU by up to 3.42% on ADE20K. Visualizations, MAD analysis, and ablation studies further validate its effectiveness.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.