Local Ai
b9558
llama.cpp release b9558 includes a Vulkan optimization that uses cm2 decode_vector for mul_mat_id B matrix loads, allowing vec4 loads and increasing BK to 64, resulting in performance speedups. The re
llama.cpp release b9558 includes a Vulkan optimization that uses cm2 decode_vector for mul_mat_id B matrix loads, allowing vec4 loads and increasing BK to 64, resulting in performance speedups. The release also includes fixes for alternative sampler name recognition (like top-k and min-p) in the common_sampler_types_from_names function to ensure compatibility with llama-server UI.
Source: llama.cpp Releases | 2026-06-08