Model Releases
gemma 4 going insane
This r/ollama thread discusses users experiencing erratic and broken behavior when running Google's Gemma 4 models locally via Ollama, including issues such as strange and unrelated responses, as well
This r/ollama thread discusses users experiencing erratic and broken behavior when running Google's Gemma 4 models locally via Ollama, including issues such as strange and unrelated responses, as well as the model skipping its thinking/reasoning output entirely . Additional reported problems include Flash Attention causing Gemma 4 31B Dense to hang indefinitely during prompt evaluation when prompts exceed approximately 3–4K tokens , and models loading into the GPU then unexpectedly jumping to CPU . These issues appear to stem from early integration bugs in Ollama's support for Gemma 4's hybrid architecture, with many users sharing workarounds and waiting on upstream fixes.
Related
- Gemma 4:e4b offloads to RAM despite having just half of VRAM used.
- Using Ollama Gemma4 models via OpenWebUI on my phone and it’s been a good experience
- Google's Gemma 4 is pretty wild. You can now run it locally with OpenClaw in 3 steps. 1. Install Ollama 2. Pull Gemma 4 model 3. Launch Open…
- Lots of love for Gemma 4! Team just told me it’s already had 10M+ downloads since last week’s launch. Gemma models have now been downloaded …
- We love seeing what you’ve built with Gemma 4, the open model family that we released last week. Here are a few fun examples, described by t…
Source: r/ollama | 2026-04-12