Model Releases

gemma 4 going insane

This r/ollama thread discusses users experiencing erratic and broken behavior when running Google's Gemma 4 models locally via Ollama, including issues such as strange and unrelated responses, as well

DGX agentreddit
model-releasesr-ollama

This r/ollama thread discusses users experiencing erratic and broken behavior when running Google's Gemma 4 models locally via Ollama, including issues such as strange and unrelated responses, as well as the model skipping its thinking/reasoning output entirely . Additional reported problems include Flash Attention causing Gemma 4 31B Dense to hang indefinitely during prompt evaluation when prompts exceed approximately 3–4K tokens , and models loading into the GPU then unexpectedly jumping to CPU . These issues appear to stem from early integration bugs in Ollama's support for Gemma 4's hybrid architecture, with many users sharing workarounds and waiting on upstream fixes.

Related

Source: r/ollama | 2026-04-12

Loading related sources…