b9112
DGX agentRelease b9112 of llama.cpp introduces a NCCL-free AllReduce implementation for LLAMA_SPLIT_MODE_TENSOR using a single-phase CUDA kernel, and adds an --allreduce flag to llama-bench to select between A
Knowledge catalogue
Release b9112 of llama.cpp introduces a NCCL-free AllReduce implementation for LLAMA_SPLIT_MODE_TENSOR using a single-phase CUDA kernel, and adds an --allreduce flag to llama-bench to select between A
This release fixes a macOS 26 target leakage issue in the v3 metallib component for the MLX framework. The fix addresses a problem where Metal library compilation was incorrectly targeting macOS 26, w