Model Releases
New model for AMD Strix Halo users: My 198B Step 3.7 Flash release was a big hit, but this one may be even better: 298B-parameter Hy3, now r…
New model for AMD Strix Halo users: My 198B Step 3.7 Flash release was a big hit, but this one may be even better: 298B-parameter Hy3, now running on a new 2-bit FPX codebook designed to map efficient
New model for AMD Strix Halo users: My 198B Step 3.7 Flash release was a big hit, but this one may be even better: 298B-parameter Hy3, now running on a new 2-bit FPX codebook designed to map efficiently into INT8 lanes on AMD hardware. The new quant: • Is 2.55% smaller than IQ2_M • Cuts end-to-end no-MTP latency by ~2.4% on a short coding prompt • Cuts latency by ~6.2% on a 19,654-token coding prompt • Stores 14.05 GB of reusable state on SSD instead of reserving scarce UMA for cache The unweighted control scored: • 81 on HermesAgent-20 • 88 on the full Tool-Eval, beating most Qwen models Performance lands around 17–25 tok/s, depending on MTP acceptance, with prefill around the 200 tok/s range. You’ll need the latest FTX inference-engine(llama fork) updates to benefit from the new optimizations. I’m also investigating KV-cache improvements that could potentially double the available context. Tencent did an excellent job with Hy3. This model even beats DeepSeek V4 Flash on several benchmarks. This may be the smartest model currently runnable on a single 128 GB system. I spent the past 24 hours validating the new quant, tool use, context handling, and Hermes performance. The fundamentals are working well, though I have barely had time to actually use the model yet. Looking forward to hearing what everyone thinks: https://huggingface.co/jcbtc/Hy3-Chadrock-FPX-IFP2-MTP
Related
- DeepSeek v4 Flash with local inference after 24h of playing with that: even with the 2 bit selective quantization GGUF, iti is the FIRST t…
- Model is available here @simonw @ivanfioravanti https://huggingface.co/mlx-community/DeepSeek-V4-Flash-2bit-DQ
- I guarantee you are sleeping on small models. Deepseek V4 Flash can do ~80% of the tasks you ask Claude or Codex for. It is 137x cheaper per…
Source: Clem Delangue (X) | 2026-07-13