b10016
[SYCL] Flash Attention with XMX engine via oneDNN (#25222) [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80
Knowledge catalogue
[SYCL] Flash Attention with XMX engine via oneDNN (#25222) [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80
B9862 is a continuous build-tagged release from llama.cpp , the open-source C/C++ LLM inference engine. Llama.cpp is the inference engine that powers most of the local-AI ecosystem, including tools li
Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other plat
What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically /v1/chat/completions streaming now matches OpenAI's wire format: rol
Learn about the productivity tool one GitHub engineer built, and how AI supported the development process. The post Build a personal organization command center with GitHub Copilot CLI appeared first
b9949 is a release build of llama.cpp, the open-source C/C++ inference engine for running large language models locally. llama.cpp releases are published frequently with build versions in the b-series
Release b9956 is a build version of llama.cpp, an open-source project for LLM inference in C/C++ . The project releases frequently, with multiple releases published in a single day , and b9956 represe
The search results do not contain specific information about v0.32.0-rc0. Based on the version numbering and context from Ollama's release patterns, v0.32.0-rc0 is a release candidate for Ollama, a la
llama.cpp b9936 is a continuous build-tagged release from the open-source llama.cpp project , which is an open-source C/C++ inference engine that powers most of the local-AI ecosystem . This release c
b9935 is a build release from the llama.cpp project, which implements LLM inference in C/C++. The project doesn't follow traditional release practices, as multiple releases can be published in a singl
The search results do not contain specific details about release b9938. However, based on the context available, b9938 is one of the regular build releases from the llama.cpp project. Llama.cpp releas
The search results show that b9946 is a specific release build tag from the llama.cpp project, though the exact details of that particular build are not visible in the page content retrieved. Based on
Based on the available information, b9905 is a build-tagged release from the llama.cpp project, which is an open-source C/C++ inference engine that powers most of the local-AI ecosystem . The project
b9908 is a build-tagged release from llama.cpp , the open-source C/C++ inference engine for large language models. llama.cpp is an open-source software library that performs inference on various large
B9932 is a continuous build-tagged release from the llama.cpp project , an open-source C/C++ inference engine for running large language models locally. llama.cpp performs inference on various large l
b9933 is a continuous build-tagged release of llama.cpp , the C/C++ inference engine for running large language models locally. This release represents an incremental update in llama.cpp's development
Based on available information, b9910 is a release tag from the llama.cpp project, an open-source C/C++ implementation for running large language model inference locally on consumer hardware. Llama.cp
B9885 is a build-tagged release of llama.cpp, an open-source C/C++ inference engine for running large language models locally. The project does not use traditional semantic versions; instead it ships
b9891 is a build release of llama.cpp, the open-source C/C++ inference engine for running large language models locally. This release enables LLM inference with minimal setup and state-of-the-art perf
v0.31.2-rc0 is a release candidate for Ollama that removes the OLLAMA_EXPERIMENT=client2 experimental flag and updates the MLX engine to the latest version . The release includes contributions from mu
Ollama v0.30.11 is a release candidate from the v0.30 series (June 2026) , which pairs the new MLX engine on Apple Silicon with continued llama.cpp improvements for enhanced compatibility across Mac a
v0.30.5 is a release candidate version of Ollama, a local large language model runner. The v0.30 series features improved compatibility and performance using llama.cpp and augments the MLX engine on A
llama.cpp is a C/C++ implementation designed to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware locally and in the cloud. Build b9515 is an interme
Ollama v0.30.2 is a patch release from the 0.30 series, which features improved compatibility and performance using llama.cpp, augmented MLX engine support on Apple Silicon, and broader model support
Ollama v0.30.4 is a patch release within the v0.30 series, which offers improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon with wider hardware support, an
The search results show general llama.cpp information and references to other recent builds (like b9484), but the specific details for b9487 were not clearly accessible. Based on the context from llam
Ollama v0.30.4 is a patch release addressing a bug in the llama-server component related to incorrect parameter wiring in the Gemma 4 model implementation. This fix ensures Gemma 4 models operate corr
Build b9483 is an intermediate release of llama.cpp, the C/C++ inference engine for running large language models locally. This build represents a specific commit snapshot from the ggml-org llama.cpp
Ollama 0.30 provides improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon for broader hardware support, and brings support for a wider range of models inclu
The search results don't contain specific information about release b9470. Based on available context, b9470 is a build release from llama.cpp, a C/C++ implementation that enables LLM inference with m
B8846 is a release version of llama.cpp, a C/C++ inference engine for running large language models locally. The release includes pre-built binaries and libraries for multiple platforms including macO
llama.cpp release **b8740** (commit `e34f042`) is a build of the open-source C/C++ LLM inference engine focused on the change 'CUDA: fuse muls' (PR #21665), which optimizes CUDA performance by fusi...
llama.cpp release **b8742** (commit `7b69125`) is a incremental build of the C/C++ LLM inference engine focused on a Vulkan backend enhancement: it adds Q1_0 quantization type support to `ggml-vulk...
llama.cpp release **b8744**, published on April 10, 2026, is a build of the ggml-org/llama.cpp C/C++ LLM inference engine. Its primary change enables the reasoning budget sampler for Gemma 4 by add...
llama.cpp release **b8747** is the latest build of the C/C++ LLM inference engine, published on April 10, 2026 (commit `fb38d6f`). Its primary change is a bug fix in the `common` layer that resolve...
**llama.cpp release b8750** is a tagged build of [llama.cpp](https://github.com/ggml-org/llama.cpp), the open-source C/C++ inference engine for large language models maintained under the ggml-org G...
llama.cpp release **b8739** is a build of the open-source C/C++ LLM inference engine that introduces HIP backend support for the CDNA4 (gfx950) GPU architecture, enabling hardware acceleration on A...
Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such
Release b9894 of llama.cpp was published on July 7, 2026 , and includes a Vulkan backend fix to check src0 type in GGML_OP_SET_ROWS to avoid failures due to unimplemented f16 support . Llama.cpp is th
The search results don't contain specific details about release b9871. Based on the pattern evident in the search results and the context, here's the summary: B9871 is a release of llama.cpp, an open-
Release b9873 is a version of llama.cpp, a C/C++ project designed to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware locally and in the cloud. The
The search results show llama.cpp releases but do not contain specific information about release b9853. Based on the context from the llama.cpp project, b9853 is likely a development build release of
Based on the available search results, the specific release details for b9857 are not fully accessible, but this entry refers to a build release from the llama.cpp project. llama.cpp is a port of Face
Release b9858 is a continuous build-tagged release of llama.cpp , the open-source C/C++ project that enables large language model inference with minimal setup and optimized performance across diverse
llama.cpp uses continuous build-tagged releases rather than traditional semantic versions. Release b9848 is a specific build from the ggml-org/llama.cpp project, which is an LLM inference implementati
B9851 is a release of llama.cpp, an LLM inference project written in C/C++ . The release represents a version update in the llama.cpp development timeline maintained on GitHub. This build tag typicall
Build b9829 is an intermediate release of llama.cpp, the open-source C/C++ project that enables users to run large language models on consumer hardware without expensive GPUs or cloud infrastructure.
B9826 is a release of llama.cpp, an LLM inference project written in C/C++ . llama.cpp enables users to run large language models on consumer hardware without expensive GPUs or cloud infrastructure .
b9811 is a release of llama.cpp , an open-source C/C++ project for running large language model inference. The release likely includes performance improvements, bug fixes, and optimizations to the cor
The search results did not provide specific information about the b9564 release. Based on the context of llama.cpp releases and the pattern observed with nearby releases (b9542, b9543, b9544, etc.), I
b9550 is a release of llama.cpp, a C/C++ implementation for LLM inference . The release represents an intermediate build in the llama.cpp development cycle, which releases frequently without tradition
llama.cpp is an open-source project that enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware . Build b9541 is an intermediate release version of llama
v0.30.7-rc1 is a release candidate for Ollama, an open-source platform for running and managing large language models locally. The v0.30 series represents improved compatibility and performance using
b9519 is a release build identifier for llama.cpp, an open-source C/C++ project that enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware . This specif
v0.30.6-rc0 is a release candidate that fixes kernel template instantiation so library symbols are exported correctly , following improvements from the v0.30 series. The v0.30 base release improved co
B9441 is a release from llama.cpp, which provides LLM inference in C/C++ . The release follows the project's versioning pattern used by the ggml-org/llama.cpp repository on GitHub, typically containin
b9401 is a release build of llama.cpp, an LLM inference framework in C/C++ that provides tools for running large language models locally. As an intermediate build in the llama.cpp project, it includes
B9403 is an intermediate build release of llama.cpp, an open-source project enabling LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware locally and in the cl
B9406 is a release build number from the llama.cpp project, a C/C++ implementation enabling LLM inference . Based on the release numbering pattern observed in the project's history, this represents an
b9412 is a release of llama.cpp, an open-source project for LLM inference in C/C++ . Build identifier b9412 represents a specific commit or version in the llama.cpp development lifecycle, following th