Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
arXiv:2607.07386v1 Announce Type: new Abstract: Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limited state size, linear attention mod