Model Releases

DeepSeek V4 paper full version is out, FP4 QAT details and stability tricks [D]

DeepSeek released the full technical report for DeepSeek-V4 on April 24, 2026, titled 'DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence.' The paper details FP4 quantization-awa

DGX agentreddit
model-releasesr-machinelearning

DeepSeek released the full technical report for DeepSeek-V4 on April 24, 2026, titled "DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence." The paper details FP4 quantization-aware training applied to MoE expert weights and the indexer QK path during training itself. Key innovations include a hybrid CSA/HCA attention scheme for ultra-long contexts, manifold-constrained hyper-connections (mHC) for stabilizing deep residual architectures, and strong engineering consolidation including the Muon optimizer at scale and deterministic kernels.

Source: r/MachineLearning | 2026-05-09

Loading related sources…