Model Releases
Would extremely high decode tok/s even be useful?
If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually
If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually useful models and fairly large models like Qwen 3.5 397B, GLM-5.2, etc Or at that speed would it just make better send to load much larger models? In which case, question still applies. E.g. Kimi K3 at the high speeds submitted by /u/LivingSwitch [link] [comments]
Related
- DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395
- Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
- Hello from 10KM high! - Thanks to Qwen 3.6 35b a3b!
Source: r/LocalLLaMA | 2026-07-30