ToolClaude Code6 recent entries5 May 2026v0.23.1Ollama v0.23.0 was released on May 3, 2026 , with v0.23.1 appearing as a patch release shortly after. The v0.23.0 release added support for Claude Desktop through Ollama Launch, with Claude Cowork and→14 May 2026v0.24.0-rc1v0.24.0-rc1 is a pre-release version focusing on improvements to Ollama server caching and the desktop launch experience, including plan-aware model gating and disabling Claude Desktop launch. The rel→22 Jul 2026v0.32.2What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @Parth
ToolOllama8 recent entries25 Jul 2026v0.32.4What's Changed x/create: quantize lm_head at 8-bit in the requested family by @jessegross in #17357 test: harden flaky updater and transfer unit tests by @dhiltgen in #17378 server: fix ps data race o→27 Jul 2026v0.32.5**Ollama – v0.32.5 Release Summary** - Version **v0.32.5** (released 27 Jul at 01:25) is the latest stable release on GitHub, with a signed commit (GPG Key ID B5690EEEBB952194). - The update includes →4 Aug 2026v0.32.6What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically /v1/chat/completions streaming now matches OpenAI's wire format: rol→10 Aug 2026v0.32.8v0.32.8 is an Oct 10, 2023 release of the ollama repository on GitHub, following a pre‑release tag v0.32.8‑rc0. The update adds Muse Glimmer support for NVIDIA, AMD and additional platforms, with the →10 Aug 2026v0.32.7Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other plat→10 Aug 2026b10342model : Granite-Switch Architecture (#25107) granite-switch: add llama.cpp backend (POC, CPU) New 'granite-switch' architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters →11 Aug 2026v0.32.9NVIDIA Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed f→11 Aug 2026v0.32.8Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such
ToolHugging Face6 recent entries10 Apr 2026b8747llama.cpp release **b8747** is the latest build of the C/C++ LLM inference engine, published on April 10, 2026 (commit `fb38d6f`). Its primary change is a bug fix in the `common` layer that resolve...→21 May 2026b9263Release b9263 of llama.cpp includes a merge of HunyuanOCR into HunyuanVL with fixes to OCR vision precision. The update consolidates OCR functionality into the HunyuanVL projector while maintaining co→2 Jun 2026v0.30.1Ollama 0.30 provides improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon for broader hardware support, and brings support for a wider range of models inclu→3 Jun 2026v0.30.4Ollama v0.30.4 is a patch release within the v0.30 series, which offers improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon with wider hardware support, an→3 Jun 2026v0.30.2Ollama v0.30.2 is a patch release from the 0.30 series, which features improved compatibility and performance using llama.cpp, augmented MLX engine support on Apple Silicon, and broader model support →6 Jun 2026v0.30.7-rc1v0.30.7-rc1 is a release candidate for Ollama, an open-source platform for running and managing large language models locally. The v0.30 series represents improved compatibility and performance using