b10036
opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (#25745) opencl: workaround for A850 compiler compat opencl: fix DX compiler version parsing and cleanup Co-
Knowledge catalogue
opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (#25745) opencl: workaround for A850 compiler compat opencl: fix DX compiler version parsing and cleanup Co-
ci : add official website link to release notes (#25728) Assisted-by: pi:llama.cpp/Qwen3.6-27B Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI en
[SYCL] Flash Attention with XMX engine via oneDNN (#25222) [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80
sycl: Increase minimum buffer size for USM system allocations (#25525) Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based on real-world experiments of overcommitting device memory
DeepseekV4: reduce graph splits (#25702) macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu
server : refactor prompt cache state ownership (#25649) server : clear checkpoints upon prompt clear server : move the prompt state data to the server_prompt_cache Assisted-by: pi:llama.cpp/Qwen3.6-27
The search did not return the specific details of the b9947 release. Based on the repository and context available: B9947 is a release of llama.cpp, which enables LLM inference in C/C++ . The release
b9948 is a build release of llama.cpp, an open-source C/C++ project for efficient large language model inference. The project ships continuous build-tagged releases using a build numbering system rath
b9949 is a release build of llama.cpp, the open-source C/C++ inference engine for running large language models locally. llama.cpp releases are published frequently with build versions in the b-series
b9950 is a release of llama.cpp, a C/C++ tool for LLM inference . This build number represents one of the frequent incremental releases from the llama.cpp project, which follows a rapid development cy
Release b9956 is a build version of llama.cpp, an open-source project for LLM inference in C/C++ . The project releases frequently, with multiple releases published in a single day , and b9956 represe
Release b9967 of llama.cpp, an LLM inference project in C/C++, includes an update to the server component allowing it to accept null sampling parameters (PR #25538). This change enables more flexible
llama.cpp is an open-source software library that performs inference on various large language models , and b9934 represents a specific build release in the project's continuous versioning system. The
b9935 is a build release from the llama.cpp project, which implements LLM inference in C/C++. The project doesn't follow traditional release practices, as multiple releases can be published in a singl
llama.cpp b9936 is a continuous build-tagged release from the open-source llama.cpp project , which is an open-source C/C++ inference engine that powers most of the local-AI ecosystem . This release c
The search results do not contain specific details about release b9938. However, based on the context available, b9938 is one of the regular build releases from the llama.cpp project. Llama.cpp releas
Release b9940 of llama.cpp was published on July 9, 2026 , and includes changes related to llama-bench initialization parameters . The release provides pre-built binaries across multiple platforms inc
b9945 is a build-tagged release of llama.cpp, an open-source library that performs inference on large language models and is co-developed alongside the GGML tensor library. The llama.cpp project does
The search results show that b9946 is a specific release build tag from the llama.cpp project, though the exact details of that particular build are not visible in the page content retrieved. Based on
Based on the available information, b9905 is a build-tagged release from the llama.cpp project, which is an open-source C/C++ inference engine that powers most of the local-AI ecosystem . The project
b9908 is a build-tagged release from llama.cpp , the open-source C/C++ inference engine for large language models. llama.cpp is an open-source software library that performs inference on various large
llama.cpp is an open-source software library that performs inference on various large language models , and the project ships continuous build-tagged releases rather than traditional semantic versioni
Based on available information, b9910 is a release tag from the llama.cpp project, an open-source C/C++ implementation for running large language model inference locally on consumer hardware. Llama.cp
b9913 is a build-tagged release from llama.cpp, an open-source software library that performs inference on various large language models. The project does not use traditional semantic versions; instea
The search results don't contain specific details about the b9914 release. Based on the context of llama.cpp releases, b9914 is a build/commit version in the llama.cpp project, an open-source tool for
b9916 is a release of llama.cpp, an open-source C/C++ implementation for LLM inference . The release represents part of the project's rapid development cycle, with binaries available for multiple plat
b9923 is a release build of llama.cpp, an open-source C/C++ project for LLM inference with minimal setup and state-of-the-art performance on various hardware . The specific b9923 build includes binary
Release b9925 of llama.cpp is a version update for the LLM inference in C/C++ project. This build is part of the project's frequent release cycle, which can publish multiple releases in a single day ,
The search results don't provide specific details about release b9929. Based on the available information and the context that llama.cpp releases frequently with tagged versions, b9929 is a specific b
Based on available information, b9931 is a release of llama.cpp, which is an open-source project for large language model inference in C/C++. As a commit-based release from the ggml-org/llama.cpp repo
B9932 is a continuous build-tagged release from the llama.cpp project , an open-source C/C++ inference engine for running large language models locally. llama.cpp performs inference on various large l
b9933 is a continuous build-tagged release of llama.cpp , the C/C++ inference engine for running large language models locally. This release represents an incremental update in llama.cpp's development
b9893 is a release build of llama.cpp with Windows OpenVINO 2026.2.1 support . The llama.cpp project uses continuous build-tagged releases rather than traditional semantic versioning , making b9893 on
Release b9894 of llama.cpp was published on July 7, 2026 , and includes a Vulkan backend fix to check src0 type in GGML_OP_SET_ROWS to avoid failures due to unimplemented f16 support . Llama.cpp is th
Release b9895 of llama.cpp includes a fix for speculative inference out-of-bounds read in ngram-map on prompt shrink, with ~2x performance gains in PP_Speed for FP32, Q4_0 and Q8_0 models. The release
Release b9902 of llama.cpp was released on July 7, 2026 , and includes support for SYCL operations including cross_entropy_loss and cross_entropy_loss_back . The release provides builds across multipl
b9879 is a release build identifier in the llama.cpp project, a C/C++ implementation of Meta's LLaMA language models. This release likely contains bug fixes, performance improvements, and feature upda
b9881 is a release of llama.cpp, an LLM inference tool written in C/C++. This release likely includes bug fixes, performance improvements, and platform support enhancements for running large language
B9884 is a llama.cpp release that addresses a Vulkan 32-bit integer overflow fix in CEIL_DIV . Released on July 6, 2026 , the build also includes platform-specific binaries for macOS, Linux, Android,
B9885 is a build-tagged release of llama.cpp, an open-source C/C++ inference engine for running large language models locally. The project does not use traditional semantic versions; instead it ships
Release b9886 of llama.cpp addresses a bug fix for K/V rotation input handling in attention mechanisms, specifically when buffers are unallocated during DFlash speculative decoding's KV-injection pass
b9891 is a build release of llama.cpp, the open-source C/C++ inference engine for running large language models locally. This release enables LLM inference with minimal setup and state-of-the-art perf
Release b9892 is a version identifier for llama.cpp, an open-source C/C++ framework for running large language model inference on consumer hardware. This specific build (b9892) represents a snapshot o
llama.cpp b9876 is a build-tagged release from the llama.cpp project , a pure C/C++ implementation of large language model inference . The project enables LLM inference with minimal setup and state-of
The search results don't contain specific details about release b9871. Based on the pattern evident in the search results and the context, here's the summary: B9871 is a release of llama.cpp, an open-
Release b9873 is a version of llama.cpp, a C/C++ project designed to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware locally and in the cloud. The
B9862 is a continuous build-tagged release from llama.cpp , the open-source C/C++ LLM inference engine. Llama.cpp is the inference engine that powers most of the local-AI ecosystem, including tools li
llama.cpp is a tool for LLM inference in C/C++ and b9864 is a release version tag from the ggml-org/llama.cpp GitHub repository. Based on the release numbering pattern and project scope, this release
b9870 is a llama.cpp release dated July 3, 2026 , which includes fixes for StepFun parser chat handling to address long reasoning loops . The release provides pre-built binaries for multiple platforms
Release b9852 is a build version of llama.cpp, a C/C++ implementation that enables large language model inference with minimal setup and state-of-the-art performance across diverse hardware platforms
The search results show llama.cpp releases but do not contain specific information about release b9853. Based on the context from the llama.cpp project, b9853 is likely a development build release of
Based on the available search results, the specific release details for b9857 are not fully accessible, but this entry refers to a build release from the llama.cpp project. llama.cpp is a port of Face
Release b9858 is a continuous build-tagged release of llama.cpp , the open-source C/C++ project that enables large language model inference with minimal setup and optimized performance across diverse
Release b9859 of llama.cpp was released on July 1, 2026 , and includes updates to allow loading precompiled binary kernels from library in the OpenCL backend . This release continues development of th
b9843 is a release of llama.cpp that reverts 'sched: reintroduce less synchronizations during split compute (#20793)' . The release includes pre-built binaries for multiple platforms including macOS,
llama.cpp is a C/C++ implementation for LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware. Release b9844 is an intermediate build of the llama.cpp project f
Release build b9846 is an intermediate build of llama.cpp, an open-source C/C++ project for LLM inference. The llama.cpp project is the main playground for developing new features for the ggml library
B9847 is a build release in the llama.cpp project, which is an open-source tool that enables LLM inference with minimal setup on a wide range of hardware . This specific build release likely contains
llama.cpp uses continuous build-tagged releases rather than traditional semantic versions. Release b9848 is a specific build from the ggml-org/llama.cpp project, which is an LLM inference implementati
b9849 is a build release from llama.cpp, which enables LLM inference in C/C++ . The release represents an intermediate build version of the llama.cpp library, a widely-used open-source project for run