Model Releases

180 tok/s generation on a 4090 with qwen 3.6. if you're on a 4090 and not running this model yet you're leaving performance on the table. 3B…

180 tok/s generation on a 4090 with qwen 3.6. if you're on a 4090 and not running this model yet you're leaving performance on the table. 3B active params at that speed is insane for agentic coding. t

DGX agentx-post
model-releasesclem-delangue--x

180 tok/s generation on a 4090 with qwen 3.6. if you're on a 4090 and not running this model yet you're leaving performance on the table. 3B active params at that speed is insane for agentic coding. thanks for the data @ErdalToprak, adding this to the comparison sheet - model: Qwen3.6-35B-A3B-UD-IQ4_XS.gguf - GPU: RTX 4090 - CUDA, f16 KV, flash attention on - n_gpu_layers=999, threads=8, batch=256, ubatch=256 - Prompt-only, 512 tokens: about 4995 tok/s - Generation-only, 128 tokens: about 180 tok/s - Mixed, 4096 prompt + 128 gen: about 2700 to…

Source: Clem Delangue (X) | 2026-04-16

Loading related sources…