Local Ai
Sliding-window beats linear attention
Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic attention with sliding window attention + attention si
Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic attention with sliding window attention + attention sinks and no post training. This could be big for memory constrained local LLM inference. EDIT: Fixed the link to the paper submitted by /u/woadwarrior [link] [comments]
Related
- Parallax: Parameterized Local Linear Attention for Language Modeling
- Are there any interesting architectural innovations that we seem to be on the verge of for LLM models or AI models that might be a big deal? (Excluding maybe N-gram, since everyone is already well aware of that one)
Source: r/LocalLLaMA | 2026-08-31