Model Releases
MoE models around A2B
There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are alre
There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are already on the heavier side if you don't have enough resources. What about the middle ground, MoE with about 2B active? I found a few, but there's very little debate about them, if any. LFM2 24B A2B (5 months old) Mellum 2 12B A2.5B (2 months old) Moondream 3.1 9B A2B (This month) VAETKI 20B A2B (7 months old) (Talk about an unknown model, it has one mention on this sub) DeepSeek V2 Lite 16B A2.4B (2024. Remember when DeepSeek was making SMALL models?) Ring Mini / Ling Mini, 16B A1.4B (2025) There are also • at least three • Nemotron fine tunes 12B A2B, and a 23B A2.8B, some 1-2 months old. Not sure what the deal is with those. Anyone uses something like this? It looks like a good size for cpu use or combined with low-end/old gpu in the 4-12GB range. In these small sizes, the increase in capability should be the most dramatic. I don't have the capacity to test properly, but hopefully some of these could beat the usual 4-9B dense suspects. Or does everyone just wanna keep simping for 1-2T models and hope something will trickle down? submitted by /u/WhoRoger [link] [comments]
Related
- Anybody else noticing how good gemma-4-26b-a4b is with one-shotting three.js?
- Hello from 10KM high! - Thanks to Qwen 3.6 35b a3b!
- 80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP
- Running Qwen 3.6 35b MoE With Zoo Code On M1 Max is Amazing! Fully local, battery-powered coding powerhouse!
Source: r/LocalLLaMA | 2026-07-23