Research
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
arXiv:2604.15769v1 Announce Type: cross Abstract: Spiking transformers achieve competitive accuracy with conventional transformers while offering 38-57imes energy efficiency on neuromorphic hardware,
arXiv:2604.15769v1 Announce Type: cross Abstract: Spiking transformers achieve competitive accuracy with conventional transformers while offering 38-57imes energy efficiency on neuromorphic hardware, yet no theoretical framework guides their design. This paper establishes the first comprehensive expressivity theory for spiking self-attention. We prove that spiking attention with Leaky Integrate-and-Fire neurons is a universal approximator of continuous permutation-equivariant functions, providing explicit spike circuit constructions including a novel lateral inhibition network for softmax normalization with proven O(1/sqrt{T}) convergence. We derive tight spike-count lower bounds via rate-distortion theory: arepsilon-approximation requires Omega(L_f^2 nd/arepsilon^2) spikes, with rigorous information-theoretic derivation. Our key insight is input-dependent bounds using measured effective dimensions (d_{ext{eff}}=47--89 for CIFAR/ImageNet), explaining why T=4 timesteps suffice despite worst-case T geq 10{,}000 predictions. We provide concrete design rules with calibrated constants (C=2.3, 95% CI: [1.9, 2.7]). Experiments on Spikformer, QKFormer, and SpikingResformer across vision and language benchmarks validate predictions with R^2=0.97 (p<0.001). Our framework provides the first principled foundation for neuromorphic transformer design.
Related
- Ge^ext{2}mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer
- Towards Green Wearable Computing: A Physics-Aware Spiking Neural Network for Energy-Efficient IMU-based Human Activity Recognition
- OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
Source: arXiv cs.AI | 2026-04-20