Model Releases
DeepSeek V4 Flash 0731 appreciation post
I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinel
I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding tasks with OpenCode? I’m genuinely amazed at what it can handle. I can throw a two-hour coding session at it, and it just keeps going until the job is done. Building integrations has never been easier - I ask OpenCode to handle it, DS tells me to hold its beer, and a little while later, it’s finished. Searching and gathering knowledge from emails? Right at your fingertips. Going through documents with Paperless NGX? No problem at all. Filling out ton of paperwork in DOCX? Easy peasy, just wrote skill in hermes, love it! OS admin work? just works! Sure, before the Q3.6 27B full FP8 on dual 3090 was really solid, but DSV4F 0731 is on a whole new level. I run a small company, and I just ordered another pair of DGX Sparks - because it genuinely feels like I now have a super capable worker on the team. I know they’re not cheap, but I’ve already saved a ton of time. I started with MiniMax M2.7 on dual Spark, and it was good - but now with DSV4F 0731? It’s just super good. And the fact that I get even better models over time, for what I already paid for, feels almost ridiculous. That’s exactly why I decided to grab another pair.. A few client tickets were literally copy-paste from the ticket system - solved, and money earned. What a time to be alive! This weekend, I’m definitely writing a ticket system integration. Can’t wait! submitted by /u/koibKop4 [link] [comments]
Related
- Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization - AI's narrative
- [[deepseek-v4-flash-0731-full-1m-context-on-a-single-rtx5090-d|[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]]]
- Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
- Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash
Source: r/LocalLLaMA | 2026-08-08