Model Releases
Congrats to @deepseek_ai team! Doing the numbers I would estimate: Pro < 14m for the final training run Flash < 4m Ratio of active params …
Congrats to @deepseek_ai team! Doing the numbers I would estimate: Pro < 14m for the final training run Flash < 4m Ratio of active params x total training tokens vs v3 Total compute costs (data prep,
Congrats to @deepseek_ai team! Doing the numbers I would estimate: Pro < 14m for the final training run Flash < 4m Ratio of active params x total training tokens vs v3 Total compute costs (data prep, tuning, testing) ~10x that Cost of taste: priceless 🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Y…
Related
- ⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par wi…
- 🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T t…
- DeepSeek V4 Pro has 1.6T total parameters, its largest model by the metric, and V4 Flash has 284B parameters; both models have a context window of 1M tokens (Vincent Chow/South China Morning Post)
- To run GLM-5.1 locally (744B params, 40B active MoE), full precision needs ~1.65TB disk + enterprise hardware like 8x H200/B200 GPUs. Minimu…
Source: Emad Mostaque (X) | 2026-04-24