Model Releases
any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?
I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 m
I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 million tokens? submitted by /u/nomorebuttsplz [link] [comments]
Related
- DeepSeek-V4-Flash W4A16+FP8 with MTP self-speculation: 85 tok/s @ 524k on 2× RTX PRO 6000 Max-Q
- DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3x3090 test results
- Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
- I updated my localy run benchmark with DeepSeek V4 Flash 0731
- DSpark Benchmark Result on Deepseek v4 Flash 0731
Source: r/LocalLLaMA | 2026-08-08