AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “llama-cpp-releases”

GridTimelineEvolution
658 results
30 Jun 2026

b9850

Local AiDGX agent

Release b9850 is a version of llama.cpp, an open-source software library for large language model inference developed alongside the GGML tensor library. This release likely includes updates, bug fixes

b9851

Local AiDGX agent

B9851 is a release of llama.cpp, an LLM inference project written in C/C++ . The release represents a version update in the llama.cpp development timeline maintained on GitHub. This build tag typicall

29 Jun 2026

b9839

Local AiDGX agent

Build b9839 of llama.cpp adds offline mode support to the llama download command for checking cached models without network access, and fixes a use-after-free bug in the URL-task callback. llama.cpp i

b9840


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai
DGX agent

b9840 is a release of llama.cpp, a tool for LLM inference in C/C++. The release enables running AI models locally on Windows, Linux, and macOS. This specific build tag likely contains bug fixes, perfo

28 Jun 2026

b9829

Local AiDGX agent

Build b9829 is an intermediate release of llama.cpp, the open-source C/C++ project that enables users to run large language models on consumer hardware without expensive GPUs or cloud infrastructure.

b9830

Local AiDGX agent

llama.cpp b9830 is a release of the open-source LLM inference project enabling local language model execution with minimal setup . The release likely includes backend improvements and optimizations fo

b9831

Local AiDGX agent

Release b9831 of llama.cpp includes backend detection improvements and synchronization enhancements, particularly for async CUDA copies and Vulkan backend operations. Llama.cpp is designed to enable L

b9832

Local AiDGX agent

B9832 is a build release of llama.cpp from the ggml-org project , which is a free and open-source tool that allows you to run AI models locally on Windows, Linux and macOS . The release likely contain

b9833

Local AiDGX agent

Release b9833 is a build of llama.cpp, an open-source project for LLM inference in C/C++. This release represents an incremental update within the llama.cpp development cycle, following the standard r

27 Jun 2026

b9822

Local AiDGX agent

b9822 is a release of llama.cpp, a project for LLM inference in C/C++ . As a build version in the ongoing development sequence of the llama.cpp project, this release likely includes bug fixes, perform

b9824

Local AiDGX agent

The search results show information about other llama.cpp releases and the project generally, but do not contain specific details about release b9824. Based on the available information about llama.cp

b9825

Local AiDGX agent

b9825 is a release of llama.cpp, the C/C++ implementation for large language model inference . This build represents an intermediate version in the ongoing development of llama.cpp, part of the ggml-o

b9826

Local AiDGX agent

B9826 is a release of llama.cpp, an LLM inference project written in C/C++ . llama.cpp enables users to run large language models on consumer hardware without expensive GPUs or cloud infrastructure .

b9827

Local AiDGX agent

Release b9827 of llama.cpp, released on June 27, 2026, added a cudaMemcpy2DAsync fast path to ggml_cuda_cpy for improved CUDA tensor copying performance. When tensors are not fully contiguous but each

b9828

Local AiDGX agent

B9828 is a llama.cpp release featuring OpenCL flash attention improvements, including reworked FA kernels for f16 and f32, prefill prepass kernels, and FA kernels for q4_0 and q8_0 quantization format

26 Jun 2026

b9810

Local AiDGX agent

b9810 is a release build of llama.cpp, an open-source software library that performs inference on various large language models such as Llama. The build identifier follows the project's versioning sch

b9811

Local AiDGX agent

b9811 is a release of llama.cpp , an open-source C/C++ project for running large language model inference. The release likely includes performance improvements, bug fixes, and optimizations to the cor

b9814

Local AiDGX agent

Based on available search results, I cannot find specific details about the b9814 release. However, b9814 is a build version identifier from the ggml-org/llama.cpp repository, which is a C/C++ impleme

b9816

Local AiDGX agent

B9816 is a release build number for llama.cpp, an open-source project that enables efficient large language model inference in C/C++ on consumer hardware. This release likely contains bug fixes, perfo

b9820

Local AiDGX agent

Release b9820 of llama.cpp introduces scheduler optimizations to reduce synchronizations during split compute and improves CUDA performance with fewer synchronizations between tokens. The update inclu

b9821

Local AiDGX agent

b9821 is the latest release of llama.cpp, a C/C++ implementation of LLM inference, which includes updates to allow the application to support --version, --licenses, and --help flags . Pre-built binari

25 Jun 2026

b9786

Local AiDGX agent

b9786 is a release of llama.cpp, a tool for LLM inference in C/C++ . Llama.cpp is a free and open-source tool that allows users to run AI models locally on Windows, Linux, and macOS . The b9786 releas

b9787

Local AiDGX agent

B9787 is a build release of llama.cpp, an open-source C/C++ implementation for LLM inference created by Georgi Gerganov. The llama.cpp project enables large language model inference with minimal setup

b9789

Local AiDGX agent

Release b9789 of llama.cpp includes a fix for quantizing mixture-of-experts models with MTP (multi-token prediction) . Binaries are provided for multiple platforms including macOS, Linux, Android, and

24 Jun 2026

b9777

Local AiDGX agent

The search results do not contain specific information about the b9777 release. Based on the available information about llama.cpp releases, b9777 is an intermediate build release from the llama.cpp p

b9784

Local AiDGX agent

Build b9784 is a release of llama.cpp, a C/C++ project that enables large language model inference with minimal setup on a wide range of hardware. The release includes pre-compiled binaries for multip

23 Jun 2026

b9769

Local AiDGX agent

The search results do not contain specific details about the b9769 release. Based on the available information about llama.cpp and its release patterns, here is a knowledge base entry: Release b9769 o

b9770

Local AiDGX agent

Release b9770 of llama.cpp addresses server functionality by fixing remote preset handling and adding tests (PR #24938). The release includes pre-built binaries for multiple platforms including macOS,

b9771

Local AiDGX agent

llama.cpp release b9771 addresses Vulkan optimization by making mul_mm ALIGNED a spec constant, reducing shader variant explosion and binary size. This release is part of the ongoing development of ll

b9773

Local AiDGX agent

Release b9773 of llama.cpp adds Vulkan support for the GET_ROWS_BACK operation . The release includes pre-built binaries for multiple platforms including macOS, Linux, Windows, and Android with variou

b9774

Local AiDGX agent

The b9774 release of llama.cpp adds Vulkan backend support for multiple operations including SQR, SQRT, SIN, COS, CLAMP, LEAKY_RELU, and NORM functions, along with fixes for non-contiguous tensor hand

b9775

Local AiDGX agent

b9775 is a release of llama.cpp published on June 23, 2026 , featuring 'server: check draft context creation error' improvements . The release includes pre-built binaries across multiple platforms inc

22 Jun 2026

b9755

Local AiDGX agent

B9755 is a release of llama.cpp, a tool for LLM inference in C/C++ . The release represents a specific build version in the active development of the project, continuing the iterative improvements to

b9756

Local AiDGX agent

Release b9756 fixes a crash in the server's edit_file function when appending at the end of a file, addressing a heap-buffer-overflow caused by improper handling of line_start -1. The fix normalizes t

b9760

Local AiDGX agent

Release b9760 of llama.cpp includes a server refactoring/generalization of the input file schema and wire-up of input_video with raw base64 support. The release includes multiple build variants across

b9761

Local AiDGX agent

The b9761 release of llama.cpp includes server improvements with model downloading moved to a dedicated process and real-time model load progress tracking via /models/sse endpoint. The release feature

b9763

Local AiDGX agent

b9763 is a release of llama.cpp, an open-source LLM inference project in C/C++ . The release includes built binaries across multiple platforms including macOS, Linux, Windows, Android, and openEuler,

21 Jun 2026

b9752

Local AiDGX agent

Release b9752 of llama.cpp focused on refactoring batch construction in the server component (PR #24843) , implementing improvements to how inference batches are handled. The release includes builds f

b9753

Local AiDGX agent

Release b9753 fixes server progress reporting for loading speculative decoding models and adds a 'stages' list feature . This update includes improvements and optimizations for the llama.cpp server co

b9754

Local AiDGX agent

Release b9754 of llama.cpp implements an AC parser for stricter grammar generation in the common/peg module , with builds available across multiple platforms including macOS, Linux, Android, and Windo

10 Jun 2026

b9589

Local AiDGX agent

llama.cpp release b9589 is a CUDA maintenance update that addresses data-race conditions in the ssm_scan_f32 kernel function by adding missing synchronization barriers for shared memory reuse. The rel

b9590

Local AiDGX agent

b9590 is a llama.cpp release that fixes the LFM2/LFM2.5 template handler which was ignoring json_schema from response_format . Released on June 10, 2026 , this build includes precompiled binaries for

b9592

Local AiDGX agent

The search results did not contain specific information about the b9592 release. Based on the available information, b9592 is a version release from the llama.cpp project, which is an LLM inference sy

9 Jun 2026

b9572

Local AiDGX agent

Release b9572 of llama.cpp fixes a bug in the ggml-cpu rms_norm_back function that produced incorrect output under in-place aliasing conditions. The release includes multiple pre-built binaries for va

b9573

Local AiDGX agent

b9573 is a build release of llama.cpp, an open-source C/C++ project for LLM inference that aims to enable language model inference with minimal setup and state-of-the-art performance on various hardwa

b9577

Local AiDGX agent

Release b9577 of llama.cpp adds a --log-prompts-dir feature to the server that writes each prompt to a separate text file in a specified directory. The release was co-authored by Xuan-Son Nguyen and i

b9578

Local AiDGX agent

Release b9578 of llama.cpp includes a refactor of video subprocess handling in the mtmd (multi-threaded multi-device) component via pull request #24316 . The release provides prebuilt binaries across

b9580

Local AiDGX agent

b9580 is a llama.cpp release that adds v_dot2_f32_f16 support in matrix-matrix multiplication and Flash Attention via Vulkan, implementing support for Valve's fp16 dot2 extension. The release also inc

b9581

Local AiDGX agent

llama.cpp b9581 is a release that includes optimization for Vulkan backend memory usage, specifically reducing iq1 shared memory usage for mul_mm operations. Released on June 9, 2026 , this build prov

b9584

Local AiDGX agent

The search results don't contain specific information about release b9584. Based on the context from other llama.cpp releases in the results, b9584 is likely an intermediate build version of llama.cpp

b9585

Local AiDGX agent

Release b9585 of llama.cpp fixes granite speech model inference by applying embedding scale when deepstack is not used . The release was published on June 9, 2026, and represents a bug fix within the

8 Jun 2026

b9554: [SYCL] Update compute runtime version to 26.x in docker (#24070)

Local AiDGX agent

This release updates the compute runtime version to 26.x within a Docker configuration for SYCL (Syclon Compute Language) support in llama.cpp. The change, addressed in pull request #24070, modernizes

b9555

Local AiDGX agent

llama.cpp is an LLM inference tool in C/C++ , and b9555 represents a specific build or version release from the ggml-org/llama.cpp GitHub repository. The project does not follow traditional release pr

b9556

Local AiDGX agent

llama.cpp b9556 release adds support for AMD RDNA3.5 graphics hardware (gfx1152 and gfx1153) in its HIP backend . The release includes compiled binaries for multiple platforms including macOS, Linux,

b9558

Local AiDGX agent

llama.cpp release b9558 includes a Vulkan optimization that uses cm2 decode_vector for mul_mat_id B matrix loads, allowing vec4 loads and increasing BK to 64, resulting in performance speedups. The re

b9561

Local AiDGX agent

B9561 is an intermediate build release of llama.cpp, the C/C++ implementation of large language model inference. Llama.cpp releases use build identifiers (b-numbers) to track development versions betw

b9563

Local AiDGX agent

b9563 is a release build of llama.cpp, an open-source software library for large language model inference that is co-developed alongside the GGML tensor library. This intermediate build includes vario

b9564

Local AiDGX agent

The search results did not provide specific information about the b9564 release. Based on the context of llama.cpp releases and the pattern observed with nearby releases (b9542, b9543, b9544, etc.), I

b9568

Local AiDGX agent

Based on available search results, I cannot find specific details about the b9568 release. B9568 is a build number from llama.cpp, an LLM inference implementation in C/C++ . This release likely contai

7 Jun 2026

b9547

Local AiDGX agent

The search did not return specific details about the b9547 release. Based on the repository context and release naming convention, b9547 is a build number for llama.cpp, an open-source C/C++ implement

← Previous
1…34567…11
Next →