Local Ai
Does Ollama Cloud prompt caching even work?
I launched a new session and tasked GLM to create an implementation plan for a spec. The plan was on the bigger side, about 5k lines. I started with 0% 5h used and ended with 80% used. A few more twea
I launched a new session and tasked GLM to create an implementation plan for a spec. The plan was on the bigger side, about 5k lines. I started with 0% 5h used and ended with 80% used. A few more tweaks were needed during review, and that's when my quota ran out. With context at 300k, it felt like every request burnt 2-3%. Small context sessions can run for a while and burn only a little bit. At this point I strongly suspect that prompt caching is not working. What are your experiences? submitted by /u/RepulsiveRaisin7 [link] [comments]
Source: r/ollama | 2026-07-24