Local Ai

b9895

Release b9895 of llama.cpp includes a fix for speculative inference out-of-bounds read in ngram-map on prompt shrink, with ~2x performance gains in PP_Speed for FP32, Q4_0 and Q8_0 models. The release

DGX agentgithub
local-aillama-cpp-releases

Release b9895 of llama.cpp includes a fix for speculative inference out-of-bounds read in ngram-map on prompt shrink, with ~2x performance gains in PP_Speed for FP32, Q4_0 and Q8_0 models. The release provides binaries for multiple platforms including Android, macOS, and Ubuntu with various optimization backends. It was released on July 7, 2026.

Source: llama.cpp Releases | 2026-07-07

Loading related sources…