Local Ai

b8860

Release b8860 of llama.cpp addresses a tensor-parallel computation issue by fixing delayed AllReduce on Gemma-4 MoE models, including optimizations to skip forward past unused nodes and allow chains o

DGX agentgithub
local-aillama-cpp-releases

Release b8860 of llama.cpp addresses a tensor-parallel computation issue by fixing delayed AllReduce on Gemma-4 MoE models, including optimizations to skip forward past unused nodes and allow chains of multiplication operations. The release includes compiled binaries for multiple platforms including macOS, Linux, Android, and Windows with support for various hardware accelerators such as Vulkan, ROCm, CUDA, and OpenVINO.

Related

Source: llama.cpp Releases | 2026-04-20

Loading related sources…