Local Ai
Taking Qwen3.5-9B quants to SOTA. New lineup incoming :)
Hey Folks, Some of you saw my Muse-Glimmer-30B SOTA line this week. I recently went all in, got a few more techniques hooked up to my pipeline, and re-quantized the 3.5 9B. And oh boy, did it demolish
Hey Folks, Some of you saw my Muse-Glimmer-30B SOTA line this week. I recently went all in, got a few more techniques hooked up to my pipeline, and re-quantized the 3.5 9B. And oh boy, did it demolish... 31 wins, 3 statistical ties, 0 losses across 34 size-matched comparisons. Every published GGUF of this model within 3% of any of my files: Unsloth's UD line, bartowski, lmstudio-community, mradermacher (i1 included), byteshape, AtomicChat. The only ties are the near-lossless Q6/Q8 tiers where everything converges. A few favorites: - My Q4_K_XL beats UD-Q4_K_XL by 23% on KLD while being smaller - and the margin holds on all six eval domains and at 32k context. - The small files are where it gets sick: my IQ2_M is 36% closer to BF16 than UD-IQ2_M at identical bytes, and scores ~10 points higher on HumanEval+ :) - MMLU, HumanEval+, MBPP+ and BFCL tool-calling (multi-turn agentic included): the flagship is statistically indistinguishable from BF16 on all of them. Full methodology on the card- eval setup, 100k-resample CIs, held-out verdict slices, the per-tier speed table, the lot. Every rival file was re-scored on the same rig against the same BF16 reference (no numbers copied from other people's cards). Happy to answer questions in the comments. Model: https://huggingface.co/AaryanK/Qwen3.5-9B-GGUF (would appreciate a like!) The point of the experiment was to test my quanting pipeline and new releases get the same treatment going forward - I've got my eye on a certain launch happening very soon 👀 (yes, the 27B...). Still a solo undergrad on rented 4090s, so the pace depends on compute money, but the lineup is real now. I'm looking for internships in AI agent orchestration and model inference. If this work looks relevant to your team: linkedin.com/in/theaaryankapoor submitted by /u/KvAk_AKPlaysYT [link] [comments]
Related
- New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)
- Local LLM Inference Optimization: The Complete Guide
- I compared all specs of the major GPUs/machines that are being used here, because bandwidth is not everything. Some of ya'll need a reality check.
Source: r/LocalLLaMA | 2026-08-14