Model Releases

DeepSeek V4 Flash on an M2 Ultra: repacked to 141 GiB losslessly, smaller than the Q4 GGUF, at 25.8 t/s (42 t/s peak)

This is one more vibe slopped custom optimization for, in this case, my hardware (m2 ultra 60 cores, 192gb). It is just a fork from llama.cpp with a few changes, it achieves: - DeepSeek V4 Flash, no k

DGX agentreddit
model-releasesr-localllama

This is one more vibe slopped custom optimization for, in this case, my hardware (m2 ultra 60 cores, 192gb). It is just a fork from llama.cpp with a few changes, it achieves: - DeepSeek V4 Flash, no kv cache quant - 141GiB model, byte-identical lossless, smaller than the public GGUFs (more room for context!) - Faster than even the M3 Ultra (16 t/s vs 25 t/s) - SSD KV cache and dynamic lanes, 1M context total, 8 lanes - PP is a bit low at ~350 t/s at 8k-32k, but SSD cache compensates for it a lot... but we could probably push this number higher, lot of compute being left on the table https://github.com/Agusx1211/llama-cpp-ds4f-m2-ultra submitted by /u/Agusx1211 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-23

Loading related sources…