Model Releases
Model is available here @simonw @ivanfioravanti https://huggingface.co/mlx-community/DeepSeek-V4-Flash-2bit-DQ
A quantized version of DeepSeek-V4-Flash has been released on Hugging Face by the MLX community in 2-bit format with dynamic quantization, making the model more efficient for inference on resource-con
A quantized version of DeepSeek-V4-Flash has been released on Hugging Face by the MLX community in 2-bit format with dynamic quantization, making the model more efficient for inference on resource-constrained devices while maintaining performance.
Source: Clem Delangue (X) | 2026-04-26