Model Releases

Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights fr…

Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B

DGX agentx-post
model-releasesemad-mostaque--x

Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B. It's capable of near Qwen3.6-35B-A3B performances, while being around a fifth of the size. This is the first model of a new series, which has been over 6 weeks of work so far fully dedicated on this. I have two models coming soon with this architecture : - a mini version of minimax m3 ( already built btw ) - a mini version of deepseek v4 flash ( already built too ). 26B and 12B. Would deffinitely love to have more compute though in order benchmark those and run more experiments and make them even better. This is I believe the fastest way for us to achieve frontier intelligence locally. @0xSero you have a lot of compute, maybe you could help me out finish my work in order to release these models opensource for everyone to use. Then I can move on to the big boys ( glm 5.2, kimi k3 and soon qwen 3.8 hehehe ). Link to the model : https://huggingface.co/Akahsizrr/fuse-1-Lite Still working on a lot of quantizations, and a checkpoint with better sparse activation.

Related

Source: Emad Mostaque (X) | 2026-08-05

Loading related sources…