AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49+ results
Model Releases

b10250

DGX agent

tests: add model resolution test on synthetic repo listings (#26172) tests: add model resolution test on synthetic repo listings Include download.cpp and arg.cpp inside a namespace with hf_cache monke

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

b10344

DGX agent

model: add MTP support for Nemotron model (#26725) model: add MTP support for Nemotron Nano model model: add mtp_flags for nemotron model address review comments Website: https://llama.app macOS/iOS:

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
model-releases
llama-cpp-releases
10 Aug 2026
Model Releases

b10144

DGX agent

server + ui: fix stream routes for model names containing a slash (#26137) server + ui: refactor resumable stream routes to query string conv_id The conversation id can embed a model name containing s

model-releasesllama-cpp-releases
27 Jul 2026
Model Releases

b10361

DGX agent

model : fix SWA not being enabled for EXAONE 4.5 (#26848) model : fix SWA not being enabled for EXAONE 4.5 load_arch_hparams tests hparams.n_layer() == 64 before LLM_KV_NEXTN_PREDICT_LAYERS has been r

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

v0.32.9

DGX agent

NVIDIA Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed f

model-releasesollama-releases
11 Aug 2026
Model Releases

b10342

DGX agent

model : Granite-Switch Architecture (#25107) granite-switch: add llama.cpp backend (POC, CPU) New 'granite-switch' architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters

model-releasesllama-cpp-releases
10 Aug 2026
Agents

Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks

DGX agent

Explore how the GitHub Copilot agentic harness delivers strong results across multiple benchmarks and leading token efficiency, while maintaining flexibility to choose among more than 20 models. The p

agentsgithub-ai-blog
25 Jun 2026
Model Releases

b10270

DGX agent

mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) convert text model main model load ok convert encoder ok speaker encoder loading ok speaker enc graph adapt vocab for backb

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

v0.32.1

DGX agent

What's Changed Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations Fixed a recurrent MLX model cache leak that could increase memory use across

model-releasesollama-releases
16 Jul 2026
Model Releases

v0.32.7

DGX agent

Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other plat

model-releasesollama-releases
10 Aug 2026
Model Releases

b10148

DGX agent

common: fix explicit -md precedence over draft sidecar resolution (#26165) common: fix explicit -md precedence over draft sidecar resolution Follow-up of #25955, an explicit --model-draft file given w

model-releasesllama-cpp-releases
27 Jul 2026
Model Releases

b10237

DGX agent

llama : MTP support for DeepSeek V3.2 (#26457) llama : MTP support for DeepSeek V3.2 model : no need to include MTP layers during DeepSeek V3.2 model type discovery Co-authored-by: Stanisław Szymczyk

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10188

DGX agent

metal: fix memory unwire if model is freed without any GPU operations (#26082) metal: fix memory leak if model is freed without any GPU operations metal: run dummy work only if residency sets are used

model-releasesllama-cpp-releases
30 Jul 2026
Model Releases

b10174

DGX agent

model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2) (#25980) model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2) Adds GLM-5.2 NextN/MTP as a --spec-type draft-m

model-releasesllama-cpp-releases
29 Jul 2026
Model Releases

b10153

DGX agent

model: Add support for Nanbeige4.2 (#25994) support nanbeige4.2 model fix fix flake8 Lint check fix loop bound check and drop redundant head_dim Co-authored-by: root lizongqiang@kanzhun.com Website: h

model-releasesllama-cpp-releases
27 Jul 2026
Model Releases

v0.30.4: llama-server: fix gemma4 patch wiring (#16477)

DGX agent

Ollama v0.30.4 is a patch release addressing a bug in the llama-server component related to incorrect parameter wiring in the Gemma 4 model implementation. This fix ensures Gemma 4 models operate corr

model-releasesollama-releases
3 Jun 2026
Local Ai

v0.20.5-rc1

DGX agent

<channel|>This release update for `local-ai` (version v0.20.5-rc1) expands model compatibility by integrating several popular large language models. It allows users to run models such as Kimi-K2.5,...

local-aiollama-releases
9 Apr 2026
Model Releases

v0.32.8

DGX agent

Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such

model-releasesollama-releases
11 Aug 2026
Model Releases

b10338

DGX agent

model-saver : fix expert shared/chunk FFN length key clobber (#26693) The saver called add_kv with LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH twice, the second time passing n_ff_chexp. gguf_set_val_u32

model-releasesllama-cpp-releases
10 Aug 2026
Model Releases

v0.32.6

DGX agent

What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically /v1/chat/completions streaming now matches OpenAI's wire format: rol

model-releasesollama-releases
4 Aug 2026
Model Releases

v0.32.0

DGX agent

What's Changed New interactive agent experience: running ollama now launches an agent to help you code and delegate work ❯ ollama Ollama 0.32.0 ▸ Chat, Code, & Work (glm-5.2:cloud) Chat with models, c

model-releasesollama-releases
14 Jul 2026
Model Releases

b10369

DGX agent

mtmd: support pocket-tts (#26871) adapt the api text model ok working impl, need verify and clean up mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no g

model-releasesllama-cpp-releases
12 Aug 2026
Model Releases

b10375

DGX agent

chat : tighten bare function parsing for Qwen models (#26793) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

model-releasesllama-cpp-releases
12 Aug 2026
Model Releases

b10312

DGX agent

server: (router) do not evict busy models (#26567) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFram

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10326

DGX agent

tts: account for the vocoder pass in the timings line (#26733) get_output runs the waveform work the pipeline defers to it, from a single trailing window to a full pass depending on the model. Measuri

model-releasesllama-cpp-releases
7 Aug 2026
Model Releases

b10295

DGX agent

model-loader : fix quantized reshaped tensor strides (#26672) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

model-releasesllama-cpp-releases
6 Aug 2026
Model Releases

b10251

DGX agent

model : support MTP in GLM-4.7-Flash (#24868) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

b10259

DGX agent

model : allow reshape of tensors during load (#26531) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCF

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

b10269

DGX agent

models : fix dflash wo_a reshape on load (#26577) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFrame

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

b10238

DGX agent

model: MTP support for Qwen3-Next (#25589) mtp for qwen3nex fix for python type-check Fix to compute num_mtp from directly mtp layer define opt_num_mtp_layers in _QwenMtpMixin and fix some comments Fi

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10242

DGX agent

CUDA: Add backend sampler for penalties sampler (#25262) sampling: enhance penalty handling in common_sampler_init Set default value for penalty_last_n based on model context if not specified. Ensure

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10244

DGX agent

model: M3: Move MSA into a new memory implementation (#26338) Move MSA logic from llama-kv-cache into llama-kv-cache-msa cont : minor cont : ws fix Co-authored-by: Georgi Gerganov ggerganov@gmail.com

model-releasesllama-cpp-releases
3 Aug 2026
Model Releases

b10225

DGX agent

model : load MiMo V2 MTP tensors only if used (#26412) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XC

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10231

DGX agent

common: support the DSpark sidecar resolution (#26458) The dspark- files resolve like the other speculative sidecars: the -hfd tag applies to them, a requested sidecar resolves without a full model at

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

b10212

DGX agent

llama : load MTP tensors only if they are really used (#26296) llama : load MTP tensors only if they are really used llama : skip loading MTP (if not used) in remaining models that support MTP Co-auth

model-releasesllama-cpp-releases
31 Jul 2026
Model Releases

b10158

DGX agent

spec: add eagle3-v3 support for gpt-oss model (#25794) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XC

model-releasesllama-cpp-releases
28 Jul 2026
Model Releases

b10173

DGX agent

model: Add Laguna-S-2.1 LLM_TYPE (#26233) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Lin

model-releasesllama-cpp-releases
28 Jul 2026
Model Releases

b10155

DGX agent

mtmd: support MiMo-V2.5 audio input (RVQ-based model) (#26190) gguf converter for mimo audio fix conv cpp impl nits nits 2 Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple

model-releasesllama-cpp-releases
27 Jul 2026
Model Releases

v0.32.3

DGX agent

What's Changed mlx update by @dhiltgen in #17332 model/parsers: finalize incomplete GLM tool calls by @dhiltgen in #17250 docs: update retirements by @mxyng in #17289 model: align Laguna with upstream

model-releasesollama-releases
23 Jul 2026
Model Releases

b10208

DGX agent

SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (#25025) SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt processing fattn-mkl: fix interleaved dst layout in nor

model-releasesllama-cpp-releases
31 Jul 2026
Model Releases

b10094

DGX agent

common: infer the speculative type from the draft repo sidecars (#25989) With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars and no --spec-type given, the draft resolved to a full mo

model-releasesllama-cpp-releases
23 Jul 2026
Local Ai

v0.32.2-rc3: test: revamp integration test entrpoints (#16560)

DGX agent

This refactors the existing integration tests into 3 priumary groups: fast, release, and library. It also refines some of the release tests to drop some of the older models and pick up newer models, w

local-aiollama-releases
21 Jul 2026
Model Releases

b10011

DGX agent

server : refactor prompt cache state ownership (#25649) server : clear checkpoints upon prompt clear server : move the prompt state data to the server_prompt_cache Assisted-by: pi:llama.cpp/Qwen3.6-27

model-releasesllama-cpp-releases
14 Jul 2026
Local Ai

b9761

DGX agent

The b9761 release of llama.cpp includes server improvements with model downloading moved to a dedicated process and real-time model load progress tracking via /models/sse endpoint. The release feature

local-aillama-cpp-releases
22 Jun 2026
Local Ai

b9544

DGX agent

B9544 is an intermediate build release of llama.cpp, an open-source C/C++ library for running large language model inference. Llama.cpp uses the GGUF model format and supports multiple hardware backen

local-aillama-cpp-releases
6 Jun 2026
Local Ai

v0.30.1

DGX agent

Ollama 0.30 provides improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon for broader hardware support, and brings support for a wider range of models inclu

local-aiollama-releases
2 Jun 2026
Model Releases

v0.30.1: llm: ignore llama-server SSE ping comments (#16443)

DGX agent

Ollama v0.30.1 addresses an issue where the LLM component now ignores Server-Sent Events (SSE) ping comments from llama-server, resolving problem #16443. This fix improves the stability and reliabilit

model-releasesollama-releases
2 Jun 2026
Local Ai

b9391

DGX agent

b9391 is a build release of llama.cpp, an open-source C/C++ inference framework for running large language models locally. llama.cpp provides lightweight, optimized model inference with support for mu

local-aillama-cpp-releases
29 May 2026
← Previous
1
Next →
525 results
← Previous
123…11
Next →