Research
Selective Rotary Position Embedding
arXiv:2511.17388v2 Announce Type: replace Abstract: Position information is essential for language modeling. In softmax transformers, Rotary Position Embeddings (extit{RoPE}) encode positions through
arXiv:2511.17388v2 Announce Type: replace Abstract: Position information is essential for language modeling. In softmax transformers, Rotary Position Embeddings (extit{RoPE}) encode positions through extit{fixed-angle} rotations, while in linear transformers, order is handled via input-dependent (selective) gating that decays past key-value associations. Selectivity has generally been shown to improve language-related tasks. Inspired by this, we introduce extit{Selective RoPE}, an extit{input-dependent} rotary embedding mechanism, that generalizes extit{RoPE}, and enables rotation in extit{arbitrary angles} for both linear and softmax transformers. We show that softmax attention already performs a hidden form of these rotations on query-key pairs, uncovering an implicit positional structure. We further show that in state-space models and gated linear transformers, the real part manages forgetting while the imaginary part encodes positions through rotations. We validate our method by equipping gated transformers with extit{Selective RoPE}, demonstrating that its input-dependent rotations improve performance in language modeling and on difficult sequence tasks like copying, state tracking, and retrieval.
Related
- Native Hybrid Attention for Efficient Sequence Modeling
- Sessa: Selective State Space Attention
- Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models
- Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
Source: arXiv cs.CL | 2026-04-27