v0.20.5-rc1
<channel|>This release update for `local-ai` (version v0.20.5-rc1) expands model compatibility by integrating several popular large language models. It allows users to run models such as Kimi-K2.5,...
Knowledge catalogue
<channel|>This release update for `local-ai` (version v0.20.5-rc1) expands model compatibility by integrating several popular large language models. It allows users to run models such as Kimi-K2.5,...
llama.cpp release **b8737** is a focused maintenance build that adds missing CUDA error handling to the ggml backend. Specifically, it checks the return values of NVIDIA CUB library calls used in t...
llama.cpp release **b8738** is a build from the [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) project introducing experimental backend-agnostic tensor parallelism, enabled via the `--...
llama.cpp release **b8739** is a build of the open-source C/C++ LLM inference engine that introduces HIP backend support for the CDNA4 (gfx950) GPU architecture, enabling hardware acceleration on A...
Ollama v0.20.5-rc0 is a release candidate for the v0.20.5 version of the Ollama open-source project, which enables users to run large language models locally. The stable v0.20.5 release that follow...
**v0.20.5-rc2** is a pre-release release candidate for the Ollama open-source local LLM runner (github.com/ollama/ollama), published around April 10, 2026, as part of the v0.20.5 release cycle. It ...
llama.cpp b9936 is a continuous build-tagged release from the open-source llama.cpp project , which is an open-source C/C++ inference engine that powers most of the local-AI ecosystem . This release c
llama.cpp is an open-source software library that performs inference on various large language models , and b9934 represents a specific build release in the project's continuous versioning system. The
b9935 is a build release from the llama.cpp project, which implements LLM inference in C/C++. The project doesn't follow traditional release practices, as multiple releases can be published in a singl
The search results do not contain specific details about release b9938. However, based on the context available, b9938 is one of the regular build releases from the llama.cpp project. Llama.cpp releas
Release b9940 of llama.cpp was published on July 9, 2026 , and includes changes related to llama-bench initialization parameters . The release provides pre-built binaries across multiple platforms inc
b9945 is a build-tagged release of llama.cpp, an open-source library that performs inference on large language models and is co-developed alongside the GGML tensor library. The llama.cpp project does
The search results show that b9946 is a specific release build tag from the llama.cpp project, though the exact details of that particular build are not visible in the page content retrieved. Based on
Based on the available information, b9905 is a build-tagged release from the llama.cpp project, which is an open-source C/C++ inference engine that powers most of the local-AI ecosystem . The project
llama.cpp is an open-source software library that performs inference on various large language models , and the project ships continuous build-tagged releases rather than traditional semantic versioni
b9908 is a build-tagged release from llama.cpp , the open-source C/C++ inference engine for large language models. llama.cpp is an open-source software library that performs inference on various large
Based on available information, b9910 is a release tag from the llama.cpp project, an open-source C/C++ implementation for running large language model inference locally on consumer hardware. Llama.cp
b9913 is a build-tagged release from llama.cpp, an open-source software library that performs inference on various large language models. The project does not use traditional semantic versions; instea
The search results don't contain specific details about the b9914 release. Based on the context of llama.cpp releases, b9914 is a build/commit version in the llama.cpp project, an open-source tool for
b9916 is a release of llama.cpp, an open-source C/C++ implementation for LLM inference . The release represents part of the project's rapid development cycle, with binaries available for multiple plat
b9923 is a release build of llama.cpp, an open-source C/C++ project for LLM inference with minimal setup and state-of-the-art performance on various hardware . The specific b9923 build includes binary
Release b9925 of llama.cpp is a version update for the LLM inference in C/C++ project. This build is part of the project's frequent release cycle, which can publish multiple releases in a single day ,
The search results don't provide specific details about release b9929. Based on the available information and the context that llama.cpp releases frequently with tagged versions, b9929 is a specific b
Based on available information, b9931 is a release of llama.cpp, which is an open-source project for large language model inference in C/C++. As a commit-based release from the ggml-org/llama.cpp repo
B9932 is a continuous build-tagged release from the llama.cpp project , an open-source C/C++ inference engine for running large language models locally. llama.cpp performs inference on various large l
b9933 is a continuous build-tagged release of llama.cpp , the C/C++ inference engine for running large language models locally. This release represents an incremental update in llama.cpp's development
B9862 is a continuous build-tagged release from llama.cpp , the open-source C/C++ LLM inference engine. Llama.cpp is the inference engine that powers most of the local-AI ecosystem, including tools li
Release b9894 of llama.cpp was published on July 7, 2026 , and includes a Vulkan backend fix to check src0 type in GGML_OP_SET_ROWS to avoid failures due to unimplemented f16 support . Llama.cpp is th
llama.cpp build **b8751** is a tagged release of the [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) project, a C/C++ library enabling high-performance LLM inference with minimal setup ...
llama.cpp release **b8752** is a tagged build of the [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) project — a C/C++ framework for efficient LLM inference on a wide range of hardware....
**llama.cpp release b8753** is a specific tagged build of the [llama.cpp](https://github.com/ggml-org/llama.cpp) project by ggml-org, an open-source C/C++ framework for running large language model...
llama.cpp release **b8754** is a tagged build of the open-source C/C++ LLM inference library maintained under the `ggml-org` GitHub organization. The release includes pre-built binaries for a wide ...
llama.cpp release **b8755** is a tagged build of the [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) project — a C/C++ library for local LLM inference. This build is one of the project'...
The search results did not return the specific changelog details for build b8756. Based on what is available and the general context of llama.cpp's rolling release model, here is a factual summary ...
llama.cpp release **b8740** (commit `e34f042`) is a build of the open-source C/C++ LLM inference engine focused on the change 'CUDA: fuse muls' (PR #21665), which optimizes CUDA performance by fusi...
llama.cpp **b8741** is an incremental build release of the open-source [llama.cpp](https://github.com/ggml-org/llama.cpp) project, which provides LLM inference in C/C++. It is one of many frequentl...
llama.cpp release **b8742** (commit `7b69125`) is a incremental build of the C/C++ LLM inference engine focused on a Vulkan backend enhancement: it adds Q1_0 quantization type support to `ggml-vulk...
llama.cpp release **b8744**, published on April 10, 2026, is a build of the ggml-org/llama.cpp C/C++ LLM inference engine. Its primary change enables the reasoning budget sampler for Gemma 4 by add...
llama.cpp release **b8746** was published on April 10, 2026 (commit `0893f50`) and consists of a single change: marking the `--split-mode tensor` option as experimental in the `--help` output (PR #...
llama.cpp release **b8747** is the latest build of the C/C++ LLM inference engine, published on April 10, 2026 (commit `fb38d6f`). Its primary change is a bug fix in the `common` layer that resolve...
The search results do not contain the specific changelog details for llama.cpp release **b8748**. The closest available data is for build b8747 (the latest at time of search), and no per-build note...
The search results do not contain specific changelog details for the exact `b8749` tag. Based on the available information about the llama.cpp project and its release cadence, here is a factual sum...
**llama.cpp release b8750** is a tagged build of [llama.cpp](https://github.com/ggml-org/llama.cpp), the open-source C/C++ inference engine for large language models maintained under the ggml-org G...
Ollama v0.20.5, released on April 9, 2026, introduces OpenClaw channel setup support, enabling users to connect WhatsApp, Telegram, Discord, and other messaging platforms via `ollama launch opencla...
Ollama v0.20.6-rc0 is a pre-release update to the Ollama local model runner, published on April 10, 2026. Key changes include adding a Hermes agent integration guide to the docs, fixing missing par...
Ollama v0.20.4 is a minor patch release published on April 7, 2026, containing two changes: improved Apple Silicon M5 performance via NAX on the MLX backend, and enabled flash attention for the Gem...
Ollama v0.20.4-rc2 is a release candidate that addresses a compatibility issue with Flash Attention (FA) for the Gemma 4 model on older GPUs. CUDA versions older than 7.5 lack the support needed t...
v0.32.8 is an Oct 10, 2023 release of the ollama repository on GitHub, following a pre‑release tag v0.32.8‑rc0. The update adds Muse Glimmer support for NVIDIA, AMD and additional platforms, with the
**Ollama – v0.32.5 Release Summary** - Version **v0.32.5** (released 27 Jul at 01:25) is the latest stable release on GitHub, with a signed commit (GPG Key ID B5690EEEBB952194). - The update includes
What's Changed x/create: quantize lm_head at 8-bit in the requested family by @jessegross in #17357 test: harden flaky updater and transfer unit tests by @dhiltgen in #17378 server: fix ps data race o
The **v0.32.2‑rc1** release (commit 4d1b53e, signed with a GPG key) introduces server-side logic to detect download stalls that occur before any data is received. It also removes the stall‑timeout set
This refactors the existing integration tests into 3 priumary groups: fast, release, and library. It also refines some of the release tests to drop some of the older models and pick up newer models, w
**Release v0.32.2‑rc0 (commit de1ce45)** adds support for Compute Capability 10.0 in the Linux “CUDA v12” preset, enabling B200‑class GPUs to use the `cuda_v12` backend even when drivers do not meet t
The search did not return the specific details of the b9947 release. Based on the repository and context available: B9947 is a release of llama.cpp, which enables LLM inference in C/C++ . The release
b9948 is a build release of llama.cpp, an open-source C/C++ project for efficient large language model inference. The project ships continuous build-tagged releases using a build numbering system rath
b9949 is a release build of llama.cpp, the open-source C/C++ inference engine for running large language models locally. llama.cpp releases are published frequently with build versions in the b-series
b9950 is a release of llama.cpp, a C/C++ tool for LLM inference . This build number represents one of the frequent incremental releases from the llama.cpp project, which follows a rapid development cy
Release b9956 is a build version of llama.cpp, an open-source project for LLM inference in C/C++ . The project releases frequently, with multiple releases published in a single day , and b9956 represe
Release b9967 of llama.cpp, an LLM inference project in C/C++, includes an update to the server component allowing it to accept null sampling parameters (PR #25538). This change enables more flexible
The search results do not contain specific information about v0.32.0-rc0. Based on the version numbering and context from Ollama's release patterns, v0.32.0-rc0 is a release candidate for Ollama, a la