Local Ai
b8860
Release b8860 of llama.cpp addresses a tensor-parallel computation issue by fixing delayed AllReduce on Gemma-4 MoE models, including optimizations to skip forward past unused nodes and allow chains o
Release b8860 of llama.cpp addresses a tensor-parallel computation issue by fixing delayed AllReduce on Gemma-4 MoE models, including optimizations to skip forward past unused nodes and allow chains of multiplication operations. The release includes compiled binaries for multiple platforms including macOS, Linux, Android, and Windows with support for various hardware accelerators such as Vulkan, ROCm, CUDA, and OpenVINO.
Related
Source: llama.cpp Releases | 2026-04-20