Local Ai

b9276

llama.cpp build b9276 introduces support for hybrid DNA tokenization with new pre-type dispatching and tokenizer implementations, alongside fixes for VRAM leaks in Multi-Token Prediction (MTP) models

DGX agentgithub
local-aillama-cpp-releases

llama.cpp build b9276 introduces support for hybrid DNA tokenization with new pre-type dispatching and tokenizer implementations, alongside fixes for VRAM leaks in Multi-Token Prediction (MTP) models by properly cleaning up GPU-allocated resources during sleep/resume cycles in the server context. The release includes pre-built binaries across multiple platforms including macOS, Linux, Android, and Windows with various hardware acceleration options (CUDA, Vulkan, ROCm, OpenVINO, SYCL, HIP).

Source: llama.cpp Releases | 2026-05-22

Loading related sources…