Model Releases

Free ~20-50% tok/s on a local llama.cpp setup if you already have a draft model sharing vocabulary with your main one. Local stack quietly g…

Free ~20-50% tok/s on a local llama.cpp setup if you already have a draft model sharing vocabulary with your main one. Local stack quietly got faster this weekend https://x.com/TechIno219886/status/20

DGX agentx-post
model-releasesclem-delangue--x

Free ~20-50% tok/s on a local llama.cpp setup if you already have a draft model sharing vocabulary with your main one. Local stack quietly got faster this weekend https://x.com/TechIno219886/status/2045940806012330239 🔴 BREAKING Speculative decoding just got real for local LLMs. llama.cpp merged speculative checkpointing (PR #19493). Expect 0-50% speedups on coding tasks with --spec-type ngram-mod. This means: • Near-instant local LLM responses • Agents that actually feel real-time • No mor…

Source: Clem Delangue (X) | 2026-04-20

Loading related sources…