Local Ai
v0.20.4
Ollama v0.20.4 is a minor patch release published on April 7, 2026, containing two changes: improved Apple Silicon M5 performance via NAX on the MLX backend, and enabled flash attention for the Gem...
Ollama v0.20.4 is a minor patch release published on April 7, 2026, containing two changes: improved Apple Silicon M5 performance via NAX on the MLX backend, and enabled flash attention for the Gemma 4 model family. It follows directly from v0.20.3 with 11 subsequent commits made to the main branch since its release.
Related
- v0.20.4-rc2: gemma4: Disable FA on older GPUs where it doesn't work (#15403)
- v0.20.5
- v0.20.5-rc1
- See you there!
- We are hosting Ollama's MLX meetup this Thursday night (April 9th) at Ollama's office in Palo Alto at 6pm. Come meet amazing people! RSVP is…
- MLX creator @awnihannun sharing the story of MLX. Apple management called him right after the launch: “why didn’t you tell us this was going…
Source: local-ai