Model Releases

Local AI just got faster. ⚡️ We worked with @ggerganov to add DFlash support in llama.cpp delivering ~2x faster inference.

Local AI just got faster. ⚡️ We worked with @ggerganov to add DFlash support in llama.cpp delivering ~2x faster inference. llama.cpp recently added DFlash support to its speculative decoding arsenal.

DGX agentx-post
model-releasesclem-delangue--x

Local AI just got faster. ⚡️ We worked with @ggerganov to add DFlash support in llama.cpp delivering ~2x faster inference. llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort! https:/…

Source: Clem Delangue (X) | 2026-07-08

Loading related sources…