Model Releases
we benchmark models nobody actually runs
qwen3.8 27B has seriously impressive benchmarks on its model card, but that's for the unquantised version. Almost everyone here will run one quant or another. Are there good benchmarks for how those q
qwen3.8 27B has seriously impressive benchmarks on its model card, but that's for the unquantised version. Almost everyone here will run one quant or another. Are there good benchmarks for how those quants perform? Kl divergence is only a rough proxy for how the distribution over tokens is preserved, but that doesn't necessarily tell you about task performance. Kl divergence could be lower due to stylistic changes, for example, that don't affect coding. This is also a general problem beyond Qwen. submitted by /u/AuspiciousApple [link] [comments]
Related
- Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
- any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?
- I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8
- Question about Quant versus Size.
Source: r/LocalLLaMA | 2026-08-17