Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
arXiv:2605.26355v1 Announce Type: cross Abstract: Standard transformer attention computes pairwise token similarity but treats all tokens as equally salient and all positions as equally local, regardl