Local Ai
Optimal 1.25 bit quantization of Qwen3.8-Flash-Next
Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if: a) it would be worth it to attempt th
Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if: a) it would be worth it to attempt this method (since they had papers detailing it) for Qwen3.8-Flash-Next b) it would be worth my time trying this as a project I wanted some feedback before I started it since I often don't know the correct scale of projects and spend too much time attempting it before I eventually realize my limits. submitted by /u/TemperatureOk3561 [link] [comments]
Source: r/LocalLLaMA | 2026-08-31