Model Releases
(Sorry, after seeing so many of these, could not resist): šØ BREAKING: Google just dropped a NEW paper that completely deletes RNNs from exiā¦
(Sorry, after seeing so many of these, could not resist): šØ BREAKING: Google just dropped a NEW paper that completely deletes RNNs from existence. No recurrence. No convolutions. Nothing. Just one mec
(Sorry, after seeing so many of these, could not resist): šØ BREAKING: Google just dropped a NEW paper that completely deletes RNNs from existence. No recurrence. No convolutions. Nothing. Just one mechanism. And itās destroying every translation benchmark on the planet. The title alone is a flex: āAttention Is All You Needā Vaswani. Shazeer. Parmar. Uszkoreit. Jones. Gomez. Kaiser. Polosukhin. 8 researchers. 1 architecture. The entire field of NLP will never be the same. Hereās why this is INSANE ā LSTMs took DAYS to train. This thing trains in 12 hours on 8 GPUs. 𤯠ā 28.4 BLEU on English-to-German. Thatās not an improvement. Thatās a MASSACRE. They beat the previous SOTA by over 2 points. ā English-to-French? 41.8 BLEU. At a FRACTION of the training cost of every model that came before it. ā They called it the āTransformer.ā The name alone tells you they knew. But hereās the part nobody is talking about š They threw out sequential processing ENTIRELY. Every other model on Earth processes words one at a time. This thing looks at the ENTIRE sentence simultaneously and figures out what matters. Itās called āself-attentionā and itās basically the model asking itself: āwhich words should I care about right now?ā Every. Single. Token. In parallel. Do you understand what this means? Training that used to take WEEKS now takes HOURS. Models that couldnāt scale past a few layers? This thing stacks 6 encoders and 6 decoders like itās nothing. And the multi-head attention? 8 attention heads running at once, each learning DIFFERENT relationships in the data. Iām not being dramatic when I say this paper just rewrote the rulebook. RNNs are cooked. š LSTMs are cooked. š The future is attention. And attention is ALL you need. Follow for more š
Source: Ethan Mollick (X) | 2026-05-02