Model Releases

The community can now download pre-quantized weights from MLX community repo on HF thanks to @LambdaAPI Model collection: https://huggingfac…

The community can now download pre-quantized weights from MLX community repo on HF thanks to @LambdaAPI Model collection: https://huggingface.co/collections/mlx-community/deepseek-v4 DeepSeek-V4-Flash

DGX agentx-post
model-releasesclem-delangue--x

The community can now download pre-quantized weights from MLX community repo on HF thanks to @LambdaAPI Model collection: https://huggingface.co/collections/mlx-community/deepseek-v4 DeepSeek-V4-Flash powering 4 parallel agents on Pi (by @badlogicgames) 🚀 Running on M3 Ultra at ~30-34 tok/s and 160-187GB peak URAM using MLX-LM. Special shoutout to @0xClandestine, @pcuenq, @kernelpool, @ivanfioravanti and others for helping optimize and shape this PR. PR: htt…

Source: Clem Delangue (X) | 2026-04-26

Loading related sources…