Local Ai

b9828

B9828 is a llama.cpp release featuring OpenCL flash attention improvements, including reworked FA kernels for f16 and f32, prefill prepass kernels, and FA kernels for q4_0 and q8_0 quantization format

DGX agentgithub
local-aillama-cpp-releases

B9828 is a llama.cpp release featuring OpenCL flash attention improvements, including reworked FA kernels for f16 and f32, prefill prepass kernels, and FA kernels for q4_0 and q8_0 quantization formats. The release also includes backend detection improvements and async CUDA copy synchronization fixes.

Source: llama.cpp Releases | 2026-06-27

Loading related sources…