Model Releases
What speeds are everyone getting with deepseek v4 flash 0731?
What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context win
What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context window of 128000, -ub/-b at 4096, “q8” unsloth’s lossless quant submitted by /u/Ambitious_Fold_2874 [link] [comments]
Related
- DeepSeek-V4-Flash-0731 unsloth gguf on A100
- Deepseek V4 Flash ~105 t/s on two Nvidia 4090d 48G (ada) in vLLM
- DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395
Source: r/LocalLLaMA | 2026-08-01