Model Releases

Deepseek V4 Flash just hit Colibri, does anyone have numbers?

I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GG

DGX agentreddit
model-releasesr-localllama

I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GGUF supported either. I'd be curious what people are getting with V100s, R9700s, etc, just to have some comparison. What's prefill like >200k context? Tg/s high enough to support agentic workloads? It's probably wishful thinking, but when I saw the release, my immediate thought was Sonnet 5 level model being "affordable" to consumers. submitted by /u/schaka [link] [comments]

Source: r/LocalLLaMA | 2026-08-05

Loading related sources…