b9505
b9505 is a latest release of llama.cpp that includes a commit adding a header to tools/server/server-http.h (#24089). The release provides precompiled binaries for multiple platforms including macOS,
Knowledge catalogue
b9505 is a latest release of llama.cpp that includes a commit adding a header to tools/server/server-http.h (#24089). The release provides precompiled binaries for multiple platforms including macOS,
Release b9509 is a build version of llama.cpp, which provides LLM inference in C/C++ . As a specific build identifier in the llama.cpp project's release history, it likely contains bug fixes, performa
b9512 is a release build of llama.cpp, which provides LLM inference in C/C++ . The release uses a version numbering system with 'b' prefixes for intermediate builds, with newer releases building upon
llama.cpp is a C/C++ implementation designed to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware locally and in the cloud. Build b9515 is an interme
llama.cpp is a C/C++ implementation that enables LLM inference with minimal setup and high performance across diverse hardware . Build b9518 is a release version from the llama.cpp project, which serv
v0.30.5 is a release candidate version of Ollama, a local large language model runner. The v0.30 series features improved compatibility and performance using llama.cpp and augments the MLX engine on A
The search results show general llama.cpp information and references to other recent builds (like b9484), but the specific details for b9487 were not clearly accessible. Based on the context from llam
Release b9488 of llama.cpp was published on June 3, 2026 , and includes support for Qwen3 SSM architectures with additions like LLM_KV_ATTENTION_RECURRENT_LAYERS . The release also fixes a bug in comm
Release b9489 of llama.cpp includes updates to hidden_act mapping in llama-model.cpp, additions of granite embedding multilingual R2 models, and support for setting hidden_activation in GGUF files. Th
b9491 is a release of llama.cpp , a C/C++ implementation of large language model inference that enables running LLMs locally with minimal dependencies. This release likely contains bug fixes, feature
B9493 is a release of llama.cpp, an LLM inference framework in C/C++ . This release includes updates to model support, such as centralized hidden activation mappings and additions for granite embeddin
The search results do not contain specific details about the b9494 release. Based on the context from the llama.cpp GitHub releases page, b9494 is an intermediate build release of llama.cpp, a C/C++ p
The search didn't return specific information about release b9495. Based on the available information about llama.cpp releases, here's a summary: llama.cpp enables LLM inference with minimal setup and
The search results do not contain specific information about release b9496. Based on the available context, b9496 is a build/release version of llama.cpp, a C/C++ implementation for efficient LLM infe
Ollama v0.30.2 is a patch release from the 0.30 series, which features improved compatibility and performance using llama.cpp, augmented MLX engine support on Apple Silicon, and broader model support
Ollama v0.30.3 adds support for the Gemma 4-12B model . This is a minor patch release that builds on the v0.30 series, which provides improved compatibility and performance improvements. The release w
Ollama v0.30.4 is a patch release within the v0.30 series, which offers improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon with wider hardware support, an
Ollama v0.30.4 is a patch release addressing a bug in the llama-server component related to incorrect parameter wiring in the Gemma 4 model implementation. This fix ensures Gemma 4 models operate corr
This release candidate fixes a Windows-specific issue where the llama-server process wasn't being properly terminated during cleanup operations. The fix addresses GitHub issue #16458 and improves the
B9466 is a build release in the llama.cpp project that includes fixes and improvements to speculative decoding functionality, specifically addressing n_outputs_max issues and extracting helper functio
The search results show recent llama.cpp releases but do not contain specific information about release b9467. Based on the context of llama.cpp releases and the GitHub repository structure, b9467 is
B9468 is an intermediate build release of llama.cpp, the open-source C/C++ library that enables large language model inference on consumer hardware. This release continues the project's rapid developm
b9469 is an intermediate build release of llama.cpp, a C/C++ implementation that enables large language model inference on consumer hardware with minimal dependencies. Build releases like b9469 repres
The search results don't contain specific information about release b9470. Based on available context, b9470 is a build release from llama.cpp, a C/C++ implementation that enables LLM inference with m
B9471 is a build release of llama.cpp, an open-source C/C++ implementation for LLM inference that enables running large language models on consumer hardware with optimized performance. This intermedia
B9473 is a release of llama.cpp, a C/C++ framework for large language model inference . The release represents a specific build commit from the ggml-org/llama.cpp repository. This entry likely documen
B9474 is a release version of llama.cpp, a tool for LLM inference in C/C++ . The release likely contains bug fixes, feature updates, and performance improvements to the llama.cpp project, which enable
b9478 is a release of llama.cpp, a project for LLM inference in C/C++ . This release represents one of the project's frequent build updates, as the llama.cpp project releases new versions regularly wi
B9480 is a release of llama.cpp, an LLM inference project in C/C++ . The release likely contains updates, bug fixes, or feature improvements to the llama.cpp codebase for running large language models
Build b9483 is an intermediate release of llama.cpp, the C/C++ inference engine for running large language models locally. This build represents a specific commit snapshot from the ggml-org llama.cpp
Ollama 0.30 provides improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon for broader hardware support, and brings support for a wider range of models inclu
Ollama v0.30.1 addresses an issue where the LLM component now ignores Server-Sent Events (SSE) ping comments from llama-server, resolving problem #16443. This fix improves the stability and reliabilit
b9451 is the latest release of llama.cpp, published on June 1, 2026 . Llama.cpp is an LLM inference project implemented in C/C++ that enables efficient local execution of large language models. This r
llama.cpp b9452 is a release of the LLM inference C/C++ project from the ggml-org repository. This release likely includes updates to the software's core functionality, bug fixes, or new features that
llama.cpp release b9453 added support for EXAONE 4.5 model implementations with vision capabilities, including markers, projector paths, and routing through a Qwen2.5-VL-style encode path with window
Release b9455 is a build of llama.cpp, a C/C++ implementation for LLM inference . This release represents a specific commit or version update to the llama.cpp project, which is used for running large
Release b9459 is a version of llama.cpp, the C/C++ implementation for LLM inference. The specific release would contain bug fixes, optimizations, or new features for the llama.cpp project, which enabl
Release b9464 is a version of llama.cpp, a project that enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware. This specific release likely contains upd
B9441 is a release from llama.cpp, which provides LLM inference in C/C++ . The release follows the project's versioning pattern used by the ggml-org/llama.cpp repository on GitHub, typically containin
b9442 is a release of llama.cpp, an open-source LLM inference tool written in C/C++. The release likely contains updates, bug fixes, or feature improvements to the llama.cpp project. For specific deta
b9444 is a release of llama.cpp that includes a server update to handle If-None-Match weak ETags . The release was published on May 31, 2026 and includes pre-built binaries for multiple platforms incl
Based on the available search results, b9430 appears to be a specific build/release tag from the llama.cpp project's development cycle. While the exact changelog for b9430 was not directly retrieved,
b9431 is a release commit of llama.cpp, an open-source C/C++ implementation for running large language model inference efficiently on consumer hardware. Based on the source material, this release like
Release b9433 is a version update from the llama.cpp project, which provides LLM inference in C/C++. Based on the GitHub releases structure, this is a specific build iteration that would contain impro
B9434 is a release build identifier from the llama.cpp project, an open-source C/C++ implementation for large language model inference. Llama.cpp releases include updates, optimizations, and bug fixes
B9436 is an intermediate build release of llama.cpp, the open-source C/C++ implementation for running large language model inference on consumer hardware. This build typically includes updates to mode
b9437 is a llama.cpp release that adds OpenCL support for bf16 format by converting to f16 . The release is part of the llama.cpp project, which is an open-source C/C++ implementation that enables loc
B9387 is a release of llama.cpp, the main playground for developing features for the GGML library, with the goal of enabling LLM inference with minimal setup and state-of-the-art performance on a wide
The search didn't return specific details about the b9388 release. Based on the available information about llama.cpp, this is likely a commit/release entry for the llama.cpp project. Here's a summary
b9391 is a build release of llama.cpp, an open-source C/C++ inference framework for running large language models locally. llama.cpp provides lightweight, optimized model inference with support for mu
B9393 is a build release of llama.cpp, the C/C++ inference framework for running large language models locally. llama.cpp enables LLM inference with minimal setup and state-of-the-art performance on a
b9394 is a release of llama.cpp, a project for LLM inference in C/C++ . The release represents a specific build or version update from the ggml-org/llama.cpp repository. Without access to the specific
b9401 is a release build of llama.cpp, an LLM inference framework in C/C++ that provides tools for running large language models locally. As an intermediate build in the llama.cpp project, it includes
B9402 is a release of llama.cpp, an open-source library that performs inference on large language models such as Llama and was developed in pure C/C++ with no dependencies. The release includes comman
B9403 is an intermediate build release of llama.cpp, an open-source project enabling LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware locally and in the cl
B9406 is a release build number from the llama.cpp project, a C/C++ implementation enabling LLM inference . Based on the release numbering pattern observed in the project's history, this represents an
b9410 is a release of llama.cpp, a C/C++ project that enables LLM inference with minimal setup and state-of-the-art performance on various hardware platforms. This build identifier represents a specif
b9411 is a release version of llama.cpp, a C/C++ implementation for LLM inference . This specific release build number represents updates to the open-source project hosted on GitHub, which provides op
b9412 is a release of llama.cpp, an open-source project for LLM inference in C/C++ . Build identifier b9412 represents a specific commit or version in the llama.cpp development lifecycle, following th
Release b9413 includes a CUDA fix that checks PTX version on the host side to guard PDL dispatch, addressing an issue where incorrect dispatching could occur on newer GPU architectures like sm_90/sm_1