Model Releases

llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the …

llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thank

DGX agentx-post
model-releasesgeorgi-gerganov--x

llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort! https://github.com/ggml-org/llama.cpp/pull/22105

Source: Georgi Gerganov (X) | 2026-07-08

Loading related sources…