Local Ai
b9789
Release b9789 of llama.cpp includes a fix for quantizing mixture-of-experts models with MTP (multi-token prediction) . Binaries are provided for multiple platforms including macOS, Linux, Android, and
Release b9789 of llama.cpp includes a fix for quantizing mixture-of-experts models with MTP (multi-token prediction) . Binaries are provided for multiple platforms including macOS, Linux, Android, and Windows with various hardware acceleration options (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL) .
Source: llama.cpp Releases | 2026-06-25