Model Releases
Qwen3.8 Flash Quants
~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. Goal: same quality band as the po
~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. Goal: same quality band as the popular Unsloth / AesSedai Q4 builds, less disk and RAM. Savings are roughly 20–30GB depending on the file you compare against. PPL is in the model card and is competitive with both. Repo: https://huggingface.co/agentionai/Qwen3.8-Flash-Next-AP-GGUF Q4 quants are the ones I would start with. Q3 and Q5 are coming. Recipe is per-layer / tailored, not a blanket lower bpw. AMD / Strix Halo: separate ROCmFP4 build that is a bit better and faster than the Q4_XS on that hardware. (https://huggingface.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF) If you try it, post your quant, RAM/VRAM, tok/s, and whether quality felt on par with Unsloth IQ4_XS / Q4_K. That is the comparison I care about. submitted by /u/Dutchnamn [link] [comments]
Related
- Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
- [[megathread-qwen38-flash-next---release-day|[Megathread] Qwen3.8-Flash-Next - Release Day]]
- Long Review: Qwen 3.8 27B is VERY good at tapping into it's real-world knowledge. It's 'overthinking' brings it to Sonnet level performance with the potential for Opus level results.
- Local agentic coding Benchmark : Qwen 3.8 27B (in many weights quants / cache quants / engine / reasoning effort) vs others.
Source: r/LocalLLaMA | 2026-08-29