Model Releases

Meituan just dropped LongCat-Flash-Lite-Sparse

It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing

DGX agentreddit
model-releasesr-localllama

It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing my Qwen 3.6 27b. submitted by /u/Gohab2001 [link] [comments]

Source: r/LocalLLaMA | 2026-07-31

Loading related sources…