Model Releases

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running_qwen3_30b_a3b_at_50_toks_on_rtx_5060_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta

DGX agentreddit
model-releasesr-localllama

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running_qwen3_30b_a3b_at_50_toks_on_rtx_5060_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta Network kernels later and here it is. It runs at 55 tok/s (61 when not recording - since the recording eats cpu and gpu capacity), significantly outperfroming llama.cpp running qwen3.5 35B at Q8 quant. Mind you - this is without MTP. MTP can speed up the generation even further, for a few cool reasons beyond the basic multiple tokens produced. Working on a blog post documenting the trick used to get this high speed up. submitted by /u/Azazelionide [link] [comments]

Source: r/LocalLLaMA | 2026-07-24

Loading related sources…