Local Ai
b9895
Release b9895 of llama.cpp includes a fix for speculative inference out-of-bounds read in ngram-map on prompt shrink, with ~2x performance gains in PP_Speed for FP32, Q4_0 and Q8_0 models. The release
Release b9895 of llama.cpp includes a fix for speculative inference out-of-bounds read in ngram-map on prompt shrink, with ~2x performance gains in PP_Speed for FP32, Q4_0 and Q8_0 models. The release provides binaries for multiple platforms including Android, macOS, and Ubuntu with various optimization backends. It was released on July 7, 2026.
Source: llama.cpp Releases | 2026-07-07