Local Ai
Request: More Transparency on Ollama Cloud Subscriptions
I've been an Ollama Cloud subscriber for ~6 months. I generally use the the latest GLM models available for coding as well as a personal instance of Open WebUI. I have tried to get answers directly vi
I've been an Ollama Cloud subscriber for ~6 months. I generally use the the latest GLM models available for coding as well as a personal instance of Open WebUI. I have tried to get answers directly via email but their support seems non-existent. I was hoping that they (or someone with more knowledge than me) might be able to clarify some issues: Lack of consistency and transparency on pricing/usage. The FAQ says limits are based on the model plus input, cached input and output tokens, and every model gets a difficulty level from 1 to 4. But there's no rate published anywhere. What does a level 3 model cost vs a level 1 model? What's the actual quota? "50x more than Free" doesn't help when free isn't a number either. OpenCode Go denominates limits in dollars. The per-request estimates people throw around are rough because they assume a token count, but I can at least see that a heavy model eats my budget faster and roughly by how much. Prompt Caching (or lack thereof) as well as the lack of communication on this topic. It's mentioned once, with no details, on the pricing FAQ page, which implies they do treat cached input differently. I can't find any evidence that it actually is (at least for billing purposes). I gave GLM-5.2 a ~200k token document and asked four follow-ups. That's roughly a million input tokens across five requests, and 800k of it was identical context I'd already sent. My quota moved as if none of it had been seen before. Also - what's the cache rate? 10% of standard input or 50% of standard input? The documentation doesn't mention it and I asked via email - no response. Happy to be wrong, but nobody will tell me either way. I will say that I do find it suspicious, or at least dubious, that Ollama doesn't return this information in their request response if they actually do discount cached input. Every other provider I've used that claims cached input rates provides this in the usage object of the response. Lack of communication and transparency (there's that word again - sensing a theme?) on new frontier models like Kimi-K3. All they said is that they are still trying to add capacity. No commitment to dates or, even more importantly, whether current subscribers will even get it. When are they just going to say that it's not available for subscribers? Just be honest. I had previously recommended Ollama Cloud subscriptions to people that had enquired but can no longer do so. submitted by /u/alexcc098 [link] [comments]
Related
- I run ollamatps.com to track TPS — after my $20 Ollama Cloud Pro plan expired, I tested the free tier and ranked all 25 models by real quota cost (same request to each)
- Not powerful but free cloud models.
- Are there truly 'free' cloud-hosted models available via Ollama for development?
Source: r/ollama | 2026-08-14