AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
527 results
Local Ai

b9391

DGX agent

b9391 is a build release of llama.cpp, an open-source C/C++ inference framework for running large language models locally. llama.cpp provides lightweight, optimized model inference with support for mu

local-aillama-cpp-releases
29 May 2026
Local Ai

b9142

DGX agent

Release b9142 is an intermediate build of llama.cpp, a C/C++ library for running large language model inference on consumer hardware. llama.cpp performs inference on various large language models and

local-aillama-cpp-releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
14 May 2026
Local Ai

v0.20.4-rc2: gemma4: Disable FA on older GPUs where it doesn't work (#15403)

DGX agent

Ollama v0.20.4-rc2 is a release candidate that addresses a compatibility issue with Flash Attention (FA) for the Gemma 4 model on older GPUs. CUDA versions older than 7.5 lack the support needed t...

local-aiollama-releases
7 Apr 2026
Model Releases

b10247

DGX agent

ggml: use dynamic allocation for split graph inputs (#22789) ggml: use dynamic allocation for split graph inputs Replace fixed-size GGML_SCHED_MAX_SPLIT_INPUTS arrays with dynamically allocated buffer

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

b10271

DGX agent

ui: CWD for agent (#26518) server : extend file_glob_search for UI pickers ui : add per-conversation working directory with picker ui : add path navigation and search scope to cwd picker Treat path-li

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

b10206

DGX agent

llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized (#25871) llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized

model-releasesllama-cpp-releases
31 Jul 2026
Model Releases

b10181

DGX agent

ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory (#26141) ggml_cuda_should_use_mmq() selects MMQ purely from the quantization type. The current MMQ configurations are designed an

model-releasesllama-cpp-releases
29 Jul 2026
Model Releases

b10172

DGX agent

ggml-webgpu: Fix some binding alias issues to support all archs, fix recurrent-state-rollback test (#25931) Add overlap glu variant to support all archs, fix recurrent-state-rollback test format Fix a

model-releasesllama-cpp-releases
28 Jul 2026
Model Releases

b10142

DGX agent

mtmd: Add Vision Support for Minimax-M3 (#25113) Add preliminary MiniMax-M3 support Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, Deep

model-releasesllama-cpp-releases
27 Jul 2026
Model Releases

v0.32.2

DGX agent

What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @Parth

model-releasesollama-releases
22 Jul 2026
Model Releases

b10016

DGX agent

[SYCL] Flash Attention with XMX engine via oneDNN (#25222) [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80

model-releasesllama-cpp-releases
15 Jul 2026
Model Releases

b10017

DGX agent

sycl: Increase minimum buffer size for USM system allocations (#25525) Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based on real-world experiments of overcommitting device memory

model-releasesllama-cpp-releases
15 Jul 2026
Model Releases

b10021

DGX agent

DeepseekV4: reduce graph splits (#25702) macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu

model-releasesllama-cpp-releases
15 Jul 2026
Local Ai

b9949

DGX agent

b9949 is a release build of llama.cpp, the open-source C/C++ inference engine for running large language models locally. llama.cpp releases are published frequently with build versions in the b-series

local-aillama-cpp-releases
10 Jul 2026
Local Ai

v0.30.7-rc1

DGX agent

v0.30.7-rc1 is a release candidate for Ollama, an open-source platform for running and managing large language models locally. The v0.30 series represents improved compatibility and performance using

local-aiollama-releases
6 Jun 2026
Local Ai

b9523

DGX agent

b9523 is a release build of llama.cpp, an open-source tool for LLM inference in C/C++ that enables language model execution with minimal setup on a wide range of hardware locally and in the cloud. The

local-aillama-cpp-releases
5 Jun 2026
Local Ai

b9493

DGX agent

B9493 is a release of llama.cpp, an LLM inference framework in C/C++ . This release includes updates to model support, such as centralized hidden activation mappings and additions for granite embeddin

local-aillama-cpp-releases
3 Jun 2026
Model Releases

v0.30.0-rc26: Merge remote-tracking branch 'upstream/main' into llama-runner-phase-0

DGX agent

This release candidate merges updates from the upstream main branch into the llama-runner-phase-0 branch, likely incorporating recent improvements and bug fixes into the development version. Version 0

model-releasesollama-releases
26 May 2026
Local Ai

b8771

DGX agent

**b8771** is a build release of [llama.cpp](https://github.com/ggml-org/llama.cpp), the open-source C/C++ library for running large language model (LLM) inference locally. Like other incremental build

local-aillama-cpp-releases
13 Apr 2026
Local Ai

v0.20.7-rc0: gemma4: add nothink renderer tests (#15554)

DGX agent

Ollama v0.20.7-rc0 is a release candidate that introduces renderer tests for the `nothink` mode specific to the Gemma 4 model (PR #15554). The `nothink` feature allows Gemma 4 to bypass its default ch

local-aiollama-releases
13 Apr 2026
Local Ai

v0.20.8-rc0: Gemma4 on MLX (#15244)

DGX agent

Ollama v0.20.8-rc0 is a release candidate that introduces MLX support for Google's Gemma 4 model family on Apple Silicon, addressing a prior limitation where Ollama would throw a `Gemma4ForConditional

local-aiollama-releases
13 Apr 2026
Local Ai

b8750

DGX agent

**llama.cpp release b8750** is a tagged build of [llama.cpp](https://github.com/ggml-org/llama.cpp), the open-source C/C++ inference engine for large language models maintained under the ggml-org G...

local-aillama-cpp-releases
10 Apr 2026
Local Ai

v0.20.6

DGX agent

Ollama v0.20.6-rc0 is a pre-release update to the Ollama local model runner, published on April 10, 2026. Key changes include adding a Hermes agent integration guide to the docs, fixing missing par...

local-aiollama-releases
10 Apr 2026
Model Releases

b10373

DGX agent

imatrix.cpp: Move finite check and only check touched experts (#26861) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS In

model-releasesllama-cpp-releases
12 Aug 2026
Model Releases

b10356

DGX agent

ci : target ROCm 7.14 for build and release (#25775) Switch ROCm from 7.2.1 to 7.14 ROCm 7.14 is the first production release using TheRock build system. It can be installed using multi-arch deliverab

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10357

DGX agent

opencl: transpose the K tile in local memory for FA prefill kernels (#26428) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED ma

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10358

DGX agent

Address review comment of PR 25532 (#26852) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework L

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10359

DGX agent

ggml-webgpu: fix CI errors from #25025 and #25262 (#26566) test new flash_attn test rebase and fix to disable subgrou matrices when max_kv_tile == 0 delete log output Add i32 support to cpy and enable

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10360

DGX agent

common/peg : suppress incomplete escape sequences (#26780) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iO

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10336

DGX agent

ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (#26134) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLE

model-releasesllama-cpp-releases
10 Aug 2026
Model Releases

b10343

DGX agent

vendor : update cpp-httplib to 0.53.0 (#26821) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramewor

model-releasesllama-cpp-releases
10 Aug 2026
Model Releases

b10353

DGX agent

ggml : require contiguous src for ROLL on CUDA and Metal (#25928) ggml_roll only asserts nb[0] == ggml_type_size, so a permuted src is a valid input, but the CUDA and Metal roll kernels index by ne al

model-releasesllama-cpp-releases
10 Aug 2026
Model Releases

b10354

DGX agent

ggml-cpu : fix CPU affinity mask being ignored on Android (#26838) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel

model-releasesllama-cpp-releases
10 Aug 2026
Model Releases

b10355

DGX agent

llama : support multi-output backend sampling (#25532) Enable backend sampling with token speculation Clamp the mask sum before converting it into the sampled index Add a numeric context parameter dec

model-releasesllama-cpp-releases
10 Aug 2026
Model Releases

b10332

DGX agent

ci: rm GGML_HIP_ROCWMMA_FATTN (#26760) Signed-off-by: Aaron Teo aaron.teo1@ibm.com Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISAB

model-releasesllama-cpp-releases
9 Aug 2026
Model Releases

b10333

DGX agent

ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (

model-releasesllama-cpp-releases
9 Aug 2026
Model Releases

b10327

DGX agent

CUDA: fix thread/block count in quantized cpy kernel launches (#26731) CUDA: fix thread/block count in quantized cpy kernel launches tests: add uneven block count cpy case Website: https://llama.app m

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10328

DGX agent

server: add initial tool isolation support (via docker) (#26507) server: add initial tool isolation support (via docker) add docs adapt get_info py: fix type check cont separate tools_io_sandbox / too

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10329

DGX agent

server, ui: only offer a working directory when a tool reads it (#26762) The working directory chip showed up as soon as the server exposed any builtin tool, so a server started with just get_datetime

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10330

DGX agent

CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767) CUDA: fuse rms_norm + mul + rope (+ view + set_rows) tests: add broadcast weight case to rms_norm_mul_rope CUDA: check memory ranges befor

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10331

DGX agent

server: report the isolate working directory from get_info (#26773) server: report the isolate working directory from get_info Without an explicit cwd, get_info fell back to the server process working

model-releasesllama-cpp-releases
8 Aug 2026
Model Releases

b10299

DGX agent

metal : avoid threadgroup matrix array instantiation in kernel_lightning_indexer (#26646) In MSL, declaring an array of matrix types like threadgroup half4x4 causes a 'no matching constructor' compila

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10301

DGX agent

cuda: fix warnings for unused variable/function (#26688) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10303

DGX agent

sycl : fix error Error OP FLASH_ATTN_EXT on arc770 (#26441) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) i

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10305

DGX agent

sycl : Support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PRE (#26568) support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PREwq update ops.md fix format issue Web

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10306

DGX agent

sycl: *glu flat path (#26354) tests: add SWIGLU perf cases perf mode had no GLU coverage. Adds SWIGLU at 17408 columns, 512 and 2048 tokens, f16 and f32, with the operands both fused and split. sycl:

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10307

DGX agent

sycl: fix UE4M3 parsing (#25608) The NVFP4 quantization format stores a scaling factor for every group of 16 weights, packed into a single UE4M3 byte. The SYCL GPU code was converting these scale valu

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10308

DGX agent

Mitigate crashing issue on Windows MSYS2 UCRT64 environment (GCC 16.1.0) (#26555) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABL

model-releasesllama-cpp-releases
7 Aug 2026
← Previous
1234…11
Next →