Model Releases
How is Deepseek v4 flash 0731 running on Ollama cloud?
I cancelled my pro plan ealier because I wanted to use new Deepseek v4 flash 0731 which was available on Openrouter through API only (not yet on ollama cloud at the time). The old Deepseek v4 flash/pr
I cancelled my pro plan ealier because I wanted to use new Deepseek v4 flash 0731 which was available on Openrouter through API only (not yet on ollama cloud at the time). The old Deepseek v4 flash/pro took away usage too fast, which is the same for GLM5.2 or Kimi K2.7 Code when the context size reach 170k+. Which I had to solve with handoff skill and new instance The new Deepseek v4 flash is insane in which I consumed 120m tokens with only 3$ spent, but that's with cache hit. To my understanding, there was no cache hit on ollama with the old version, that's why the usage basically evaporated when using deepseek v4 I could comeback to ollama cloud if DS4F 0731 is working properly now... https://preview.redd.it/lwwe2fi1kphh1.png?width=1281&format=png&auto=webp&s=51b8f3ff2bb8cd917b90b4c79c1e9130bc6abb0a submitted by /u/quantanhoi [link] [comments]
Source: r/ollama | 2026-08-06