Model Releases
DeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (and back again!)
Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but.. really not that impressed with GLM 5.3 - overly v
Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but.. really not that impressed with GLM 5.3 - overly verbose and takes for ever (was getting around 22 tok/s on dual spark setup). Have now got myself setup as DeepSeek V4 0731 running on 2 of the sparks and Qwen 3.8 Flash running on the other two. DS is my plan and build and Qwen is explore / scout / subagent work. Seems to be running as a pretty good setup. Anyone else tried out GLM 5.3 Flash on DGX Sparks yet? What's your thoughts on the new GLM and Qwen models? submitted by /u/Legitimate_Hat_7852 [link] [comments]
Related
- DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration
- Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
- Qwen 3.8 27b vs Deepseek Flash
- Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?
Source: r/LocalLLaMA | 2026-08-27