Local Ai
b9105
Release b9105 of llama.cpp includes updates to the AllReduce implementation for CUDA, introducing a NCCL-free provider for tensor parallelism that pipelines data transfers and GPU reduction operations
Release b9105 of llama.cpp includes updates to the AllReduce implementation for CUDA, introducing a NCCL-free provider for tensor parallelism that pipelines data transfers and GPU reduction operations. The release also features renaming of llama-bench flags and fixes to environment variable handling for provider selection.
Source: llama.cpp Releases | 2026-05-11