Local Ai
b9828
B9828 is a llama.cpp release featuring OpenCL flash attention improvements, including reworked FA kernels for f16 and f32, prefill prepass kernels, and FA kernels for q4_0 and q8_0 quantization format
B9828 is a llama.cpp release featuring OpenCL flash attention improvements, including reworked FA kernels for f16 and f32, prefill prepass kernels, and FA kernels for q4_0 and q8_0 quantization formats. The release also includes backend detection improvements and async CUDA copy synchronization fixes.
Source: llama.cpp Releases | 2026-06-27