Local Ai
b9119
Build b9119 of llama.cpp introduces CUDA optimizations including an NCCL-free AllReduce implementation for tensor-parallel inference and updates to the llama-bench tool for managing reduction provider
Build b9119 of llama.cpp introduces CUDA optimizations including an NCCL-free AllReduce implementation for tensor-parallel inference and updates to the llama-bench tool for managing reduction providers. The release improves logging by ensuring WARN and ERROR level messages remain visible in non-verbose mode to surface important issues like unavailable reduction providers.
Source: llama.cpp Releases | 2026-05-12