Local Ai
b9122
Release b9122 of llama.cpp addresses precision issues for multimodal models through ggml-webgpu, including fixes for GELU functions, flash attention tile implementations, and type conflict resolution.
Release b9122 of llama.cpp addresses precision issues for multimodal models through ggml-webgpu, including fixes for GELU functions, flash attention tile implementations, and type conflict resolution. The update includes corrections to avoid NaN values during computation and adjustments to safer numerical ranges for exponential operations.
Source: llama.cpp Releases | 2026-05-12