Local Ai

b9558

llama.cpp release b9558 includes a Vulkan optimization that uses cm2 decode_vector for mul_mat_id B matrix loads, allowing vec4 loads and increasing BK to 64, resulting in performance speedups. The re

DGX agentgithub
local-aillama-cpp-releases

llama.cpp release b9558 includes a Vulkan optimization that uses cm2 decode_vector for mul_mat_id B matrix loads, allowing vec4 loads and increasing BK to 64, resulting in performance speedups. The release also includes fixes for alternative sampler name recognition (like top-k and min-p) in the common_sampler_types_from_names function to ensure compatibility with llama-server UI.

Source: llama.cpp Releases | 2026-06-08

Loading related sources…