Local Ai
b9109
Release b9109 of llama.cpp introduces refinements to the CUDA tensor parallelism AllReduce implementation, including renaming the --allreduce flag to --reduction-provider and updates to NCCL-free AllR
Release b9109 of llama.cpp introduces refinements to the CUDA tensor parallelism AllReduce implementation, including renaming the --allreduce flag to --reduction-provider and updates to NCCL-free AllReduce functionality for multi-GPU support. The release also improves logging by ensuring WARN and ERROR messages are visible in non-verbose mode to flag issues like unavailable reduction providers.
Source: llama.cpp Releases | 2026-05-11