Model Releases
Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights fr…
Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B
Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B. It's capable of near Qwen3.6-35B-A3B performances, while being around a fifth of the size. This is the first model of a new series, which has been over 6 weeks of work so far fully dedicated on this. I have two models coming soon with this architecture : - a mini version of minimax m3 ( already built btw ) - a mini version of deepseek v4 flash ( already built too ). 26B and 12B. Would deffinitely love to have more compute though in order benchmark those and run more experiments and make them even better. This is I believe the fastest way for us to achieve frontier intelligence locally. @0xSero you have a lot of compute, maybe you could help me out finish my work in order to release these models opensource for everyone to use. Then I can move on to the big boys ( glm 5.2, kimi k3 and soon qwen 3.8 hehehe ). Link to the model : https://huggingface.co/Akahsizrr/fuse-1-Lite Still working on a lot of quantizations, and a checkpoint with better sparse activation.
Related
- Today we’re announcing Ternary Bonsai: Top intelligence at 1.58 bits Using ternary weights {-1, 0, +1}, we built a family of models that are…
- Built and released BetterGPT-150M – A compact 150M parameter completion model (+ live HF Space demo)
Source: Emad Mostaque (X) | 2026-08-05