Local Ai
b9580
b9580 is a llama.cpp release that adds v_dot2_f32_f16 support in matrix-matrix multiplication and Flash Attention via Vulkan, implementing support for Valve's fp16 dot2 extension. The release also inc
b9580 is a llama.cpp release that adds v_dot2_f32_f16 support in matrix-matrix multiplication and Flash Attention via Vulkan, implementing support for Valve's fp16 dot2 extension. The release also includes updates to the RPC protocol patch version for a new GGML_OP_COL2IM_1D operation.
Source: llama.cpp Releases | 2026-06-09