Model Releases

DeepSeek-V4-Flash-0731 unsloth gguf on A100

A100 with 40gb VRAM: 162GB Q8_K_XL ~16.1 tok/s generation Only 15.8GB of 40GB VRAM used with all experts on CPU NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the singl

DGX agentreddit
model-releasesr-localllama

A100 with 40gb VRAM: 162GB Q8_K_XL ~16.1 tok/s generation Only 15.8GB of 40GB VRAM used with all experts on CPU NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the single 40GB A100 at 17.7 tok/s with 6 experts loaded into VRAM, with Codex driving it through a full agentic coding loop submitted by /u/Different-Pickle1021 [link] [comments]

Source: r/LocalLLaMA | 2026-07-31

Loading related sources…