Local Ai

Ollama cloud + GLM 5.1 slow and stupid or am I?

This Reddit thread likely discusses user frustrations with performance issues when running GLM-5.1 via Ollama's cloud inference option (`glm-5.1:cloud`), a flagship agentic coding model from Z.AI. Com

DGX agentreddit
local-air-ollama

This Reddit thread likely discusses user frustrations with performance issues when running GLM-5.1 via Ollama's cloud inference option (glm-5.1:cloud), a flagship agentic coding model from Z.AI. Community reports and GitHub issues confirm that Ollama Cloud has experienced significant reliability problems, including both /api/chat and /api/generate endpoints returning empty responses or timeouts across all cloud models, not limited to GLM-5.1 specifically. Users accessing GLM-5.1 through Ollama's cloud path rather than self-hosting may encounter slow or degraded responses that could be mistaken for poor model quality, when in fact GLM-5.1 is designed as a next-generation flagship model for agentic engineering with strong coding capabilities, achieving state-of-the-art performance on benchmarks like SWE-Bench Pro.

Related

Source: r/ollama | 2026-04-15

Loading related sources…