TechniqueRLHF / Alignment3 recent entries10 Apr 2026v0.20.6Ollama v0.20.6-rc0 is a pre-release update to the Ollama local model runner, published on April 10, 2026. Key changes include adding a Hermes agent integration guide to the docs, fixing missing par...→30 Apr 2026v0.22.1-rc1v0.22.1-rc1 is a pre-release version of Ollama that includes improvements to MLX sampler batching, tokenizer BPE offset handling, NVIDIA TensorRT Model Optimizer support, and fixes for desktop app sta→23 Jul 2026v0.32.3What's Changed mlx update by @dhiltgen in #17332 model/parsers: finalize incomplete GLM tool calls by @dhiltgen in #17250 docs: update retirements by @mxyng in #17289 model: align Laguna with upstream
TechniqueAgents8 recent entries7 Jul 2026v0.31.2Ollama v0.31.2-rc1 is a pre-release version released on July 6, 2026. This release includes CI improvements to avoid unbounded parallelism, fixes for CUDA toolkit lookup, updates to cloud documentatio→14 Jul 2026v0.32.0What's Changed New interactive agent experience: running ollama now launches an agent to help you code and delegate work ❯ ollama Ollama 0.32.0 ▸ Chat, Code, & Work (glm-5.2:cloud) Chat with models, c→16 Jul 2026v0.32.1What's Changed Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations Fixed a recurrent MLX model cache leak that could increase memory use across →22 Jul 2026v0.32.2What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @Parth→25 Jul 2026v0.32.4What's Changed x/create: quantize lm_head at 8-bit in the requested family by @jessegross in #17357 test: harden flaky updater and transfer unit tests by @dhiltgen in #17378 server: fix ps data race o→10 Aug 2026v0.32.7Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other plat→11 Aug 2026v0.32.9NVIDIA Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed f→11 Aug 2026v0.32.8Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such
TechniqueFine-tuning2 recent entries2 Jun 2026v0.30.1Ollama 0.30 provides improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon for broader hardware support, and brings support for a wider range of models inclu→3 Jun 2026v0.30.4Ollama v0.30.4 is a patch release within the v0.30 series, which offers improved compatibility and performance using llama.cpp, augments the MLX engine on Apple Silicon with wider hardware support, an
TechniqueMultimodal2 recent entries7 Jul 2026v0.31.2-rc2: llm: allow iGPU mmproj offload with fit padding (#16996)This release candidate introduces support for offloading the image GPU projection (mmproj) to an integrated GPU when using fit padding in Ollama's LLM processing, addressing technical improvements for→10 Aug 2026v0.32.7Muse Glimmer Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other plat