Model Releases

Congrats to @deepseek_ai team! Doing the numbers I would estimate: Pro < 14m for the final training run Flash < 4m Ratio of active params …

Congrats to @deepseek_ai team! Doing the numbers I would estimate: Pro < 14m for the final training run Flash < 4m Ratio of active params x total training tokens vs v3 Total compute costs (data prep,

DGX agentx-post
model-releasesemad-mostaque--x

Congrats to @deepseek_ai team! Doing the numbers I would estimate: Pro < 14m for the final training run Flash < 4m Ratio of active params x total training tokens vs v3 Total compute costs (data prep, tuning, testing) ~10x that Cost of taste: priceless 🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Y…

Related

Source: Emad Mostaque (X) | 2026-04-24

Loading related sources…