Research

Gated DeltaNet has been one of my favorite 'hybrid attention' newcomers in the good old transformer stack. Excited to see Gated DeltaNet-2. …

Gated DeltaNet has been one of my favorite 'hybrid attention' newcomers in the good old transformer stack. Excited to see Gated DeltaNet-2. Adding it to my reading stack. In the meantime, I have a pri

DGX agentx-post
researchsebastian-raschka--x

Gated DeltaNet has been one of my favorite "hybrid attention" newcomers in the good old transformer stack. Excited to see Gated DeltaNet-2. Adding it to my reading stack. In the meantime, I have a primer on Gated DeltaNet here: https://magazine.sebastianraschka.com/i/177848019/26-gated-deltanet Gated DeltaNet-2 is here. 🚀 🔥 New paper: Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention Gated DeltaNet-2 outperforms KDA and Mamba-3, the latest and best recurrent architectures, head to head at 1.3B. 🏆 💡 Here's the idea behind it: Linear attention squeezes …

Source: Sebastian Raschka (X) | 2026-05-21

Loading related sources…