Model Releases

Probably the best way to run DS4 flash on a mac right now (192gb+ vram)

Found this quant, so thought I would share, since its the best I've found so far for running on my mac (m3 ultra). It's got dspark/mtp support so runs faster than anything else I've tried. The tok/s o

DGX agentreddit
model-releasesr-localllama

Found this quant, so thought I would share, since its the best I've found so far for running on my mac (m3 ultra). It's got dspark/mtp support so runs faster than anything else I've tried. The tok/s on this code run actually increased as generation went on, started at 34tok/s, ended at 43tok/s. The cached tokens were the default chat prompt, and the 13k was the query I sent. https://huggingface.co/Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX submitted by /u/Professional-Bear857 [link] [comments]

Source: r/LocalLLaMA | 2026-08-04

Loading related sources…