Model Releases

KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates

Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5

DGX agentreddit
model-releasesr-localllama

Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5_K_S 64k context, Gemma 4 31B Q5_K_S 16k context Standard quants, extended: q6_0 and q6_1, and low-bit types from q2_0 to q3_1 KVarN: Variance-Normalized KV-Cache by Huawei, implemented in BeeLlama Precision Tail: keeping latest X tokens of KV cache in (B)F16, implemented in BeeLlama 413 configurations in total: 238 with Qwen 3.6 27B, 175 with Gemma 4 31B The Recommendation Ladder Full benchmark results, setup, method, analysis, explanations and everything else can be found in the article. 1. Qwen Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for bf16 0 4096.00 0 0.00005 Reference q8_0 1024 2272.00 0.000897 0.087699 Standard fidelity with a precision tail kvarn8 1024 2256.00 0.000871 0.087639 Best measured quality below BF16 q8_0 0 2176.00 0.000909 0.093029 Standard fidelity q8_0-q6_0 1024 2016.00 0.000894 0.091098 q8_0 quality within noise, 256.00 MiB less kvarn6 1024 1744.00 0.000879 0.084629 The high-end value pick kvarn6-kvarn5 1024 1616.00 0.000886 0.092778 Much cheaper, almost as good kvarn5 1024 1488.00 0.000897 0.087666 Highest value in mid-range q5_0-q4_1 1024 1440.00 0.000966 0.089128 Standard when VRAM-constrained kvarn5-kvarn4 1024 1360.00 0.000936 0.089469 Balanced default q4_0 1024 1248.00 0.001057 0.104486 Compact standard kvarn4 1024 1232.00 0.000994 0.090391 Cleaner than q4_0 for less memory kvarn4-kvarn3 1024 1104.00 0.001112 0.113968 Smallest recommended tier kvarn3 1024 976.00 0.001316 0.139558 When the context must fit kvarn3-kvarn2 1024 848.00 0.002424 0.23878 Emergency compression kvarn2 1024 720.00 0.003811 0.450496 Last resort 2. Qwen Standard-Only Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for bf16 0 4096.00 0 0.00005 Reference q8_0 0 2176.00 0.000909 0.093029 Compression with minimal losses q8_0-q6_0 0 1920.00 0.000937 0.093575 256.00 MiB below q8_0 q6_0 0 1664.00 0.00096 0.091134 The high-end value pick q6_0-q5_0 0 1536.00 0.001054 0.09467 Balanced default q5_0 0 1408.00 0.001154 0.09707 Last tier before the cliff q5_0-q4_1 0 1344.00 0.001433 0.122096 Default when VRAM-constrained q5_0-q4_0 0 1280.00 0.001516 0.121068 64.00 MiB cheaper, worse median q4_0 0 1152.00 0.001846 0.154408 Smallest recommended tier q4_0-q3_0 0 1024.00 0.003313 0.218912 When the context must fit q3_0 0 896.00 0.004696 0.304186 Emergency compression q2_0 0 640.00 0.019374 1.198902 Last resort 3. Gemma Cache Tail KV cache (MiB) Median KLD 99.9% KLD What it is for bf16 0 2480.00 0 0.000047 Reference q8_0 0 1317.50 0.0371 16.813929 General default at full prefill speed q8_0-q6_0 0 1162.50 0.040875 16.839821 155.00 MiB below q8_0 q6_0 0 1007.50 0.042636 17.30599 Last tier before the cliff q6_0-q5_0 0 930.00 0.055236 17.26157 Stronger K side, 77.50 MiB above q5_0 q5_0 0 852.50 0.061747 18.731647 Memory floor for usable quality q5_0-q4_0 0 775.00 0.109427 19.183374 Asymmetric compact q4_0 0 697.50 0.134091 20.442234 Budget body before the huge cliff q4_0-q3_0 0 620.00 0.381216 22.304634 When the context must fit q3_0 0 542.50 0.504075 23.15744 Emergency compression q2_0 0 387.50 2.95758 27.834961 Last resort submitted by /u/Anbeeld [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-06

Loading related sources…