Model Releases

A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing

Throughput observation more than a demo. Box and recording are sudoingX's on X, shared with his okay; I work on Ling at inclusionAI. He handed the web UI on his llama-server a 33-token prompt — build

DGX agentreddit
model-releasesr-localllama

Throughput observation more than a demo. Box and recording are sudoingX's on X, shared with his okay; I work on Ling at inclusionAI. He handed the web UI on his llama-server a 33-token prompt — build a gpu monitoring dashboard frontend, dummy data, premium design — and left it running. Ling-3.0-flash on the community Q5 GGUF, one Spark. Nothing else on the box but Xorg. Single response, no turns: eval time = 424035.62 ms / 15128 tokens (28.03 ms per token, 35.68 tokens per second) truncated = 0 The total isn't the interesting bit. At n_decoded 2793 the log says 35.62 t/s. At 15062 it says 35.68. Twelve thousand more tokens of KV cache and decode sat still. What came out is a dashboard frontend on simulated data — Math.random() drift and two GPUs that box doesn't have. That's what he asked for so it isn't a miss, but it is not reading the GPU, and the word dummy is right there in the prompt on screen. Seven minutes of generation. I don't have a coherence check on the output past the fact that it renders. submitted by /u/AcanthisittaOk1699 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-13

Loading related sources…