Model Releases

Would extremely high decode tok/s even be useful?

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually

DGX agentreddit
model-releasesr-localllama

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually useful models and fairly large models like Qwen 3.5 397B, GLM-5.2, etc Or at that speed would it just make better send to load much larger models? In which case, question still applies. E.g. Kimi K3 at the high speeds submitted by /u/LivingSwitch [link] [comments]

Related

Source: r/LocalLLaMA | 2026-07-30

Loading related sources…