Model Releases
A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing
Throughput observation more than a demo. Box and recording are sudoingX's on X, shared with his okay; I work on Ling at inclusionAI. He handed the web UI on his llama-server a 33-token prompt — build
Throughput observation more than a demo. Box and recording are sudoingX's on X, shared with his okay; I work on Ling at inclusionAI. He handed the web UI on his llama-server a 33-token prompt — build a gpu monitoring dashboard frontend, dummy data, premium design — and left it running. Ling-3.0-flash on the community Q5 GGUF, one Spark. Nothing else on the box but Xorg. Single response, no turns: eval time = 424035.62 ms / 15128 tokens (28.03 ms per token, 35.68 tokens per second) truncated = 0 The total isn't the interesting bit. At n_decoded 2793 the log says 35.62 t/s. At 15062 it says 35.68. Twelve thousand more tokens of KV cache and decode sat still. What came out is a dashboard frontend on simulated data — Math.random() drift and two GPUs that box doesn't have. That's what he asked for so it isn't a miss, but it is not reading the GPU, and the word dummy is right there in the prompt on screen. Seven minutes of generation. I don't have a coherence check on the output past the fact that it renders. submitted by /u/AcanthisittaOk1699 [link] [comments]
Related
- Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
- DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000
- Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash
Source: r/LocalLLaMA | 2026-08-13