Research
Transformer Approximations from ReLUs
arXiv:2604.24878v1 Announce Type: new Abstract: We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation t
arXiv:2604.24878v1 Announce Type: new Abstract: We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation targets. Importantly, it yields target-specific, economic resource bounds beyond universal approximation statements. We showcase the recipe on multiplication, reciprocal computation, and min/max primitives. These results provide new analytical tools for analyzing softmax transformer models.
Related
- Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
- Higher Order Approximation Rates for ReLU CNNs in Korobov Spaces
- Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
- Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
- Gating Enables Curvature: A Geometric Expressivity Gap in Attention
- Explicit integral representations and quantitative bounds for two-layer ReLU networks
Source: arXiv cs.LG | 2026-04-29