Research
Semidirect Fourier Delta Attention: Phase-Controlled Delta Memory with Constructive Chunk-WY Kernels
arXiv:2607.11897v1 Announce Type: new Abstract: Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-c
arXiv:2607.11897v1 Announce Type: new Abstract: Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-context memory. We introduce Semidirect Fourier Delta Attention (SFDA), a phase-controlled generalization of Kimi Delta Attention that replaces real diagonal decay with block-rotational Fourier control: [ S_t=(I-eta_t k_tk_t^)Lambda_tS_{t-1}+eta_tk_tv_t^, qquad Lambda_t=iag(alpha_todot e^{iheta_t}). ] Our main result is a constructive chunk-WY factorization for products (A_t=Lambda_t-u_tr_t^), giving [ A_tdots A_1=Gamma_t-Y_tM_tW_t^ ] with rank growth bounded inside fixed chunks. This yields an exact affine chunk transfer, formal stability and complexity bounds, and a compact characterization of phase-plus-low-rank memory. We verify the algebra numerically and show in toy state-tracking experiments that SFDA learns cyclic memory where the phase-disabled KDA baseline remains near chance. Fused kernels and large-scale language-model comparisons are left to future work.
Related
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
- A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets
- Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference
- OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
Source: arXiv cs.LG | 2026-07-15