Local Ai

Does Ollama Cloud prompt caching even work?

I launched a new session and tasked GLM to create an implementation plan for a spec. The plan was on the bigger side, about 5k lines. I started with 0% 5h used and ended with 80% used. A few more twea

DGX agentreddit
local-air-ollama

I launched a new session and tasked GLM to create an implementation plan for a spec. The plan was on the bigger side, about 5k lines. I started with 0% 5h used and ended with 80% used. A few more tweaks were needed during review, and that's when my quota ran out. With context at 300k, it felt like every request burnt 2-3%. Small context sessions can run for a while and burn only a little bit. At this point I strongly suspect that prompt caching is not working. What are your experiences? submitted by /u/RepulsiveRaisin7 [link] [comments]

Source: r/ollama | 2026-07-24

Loading related sources…