Model Releases
Getting the most out of MTP
If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving
If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving a lot of performance on the table if you don't tune n_max. Can be easily missing out on 50-100% of the possible performance on some models. I ran some benchmarks against the various models I am using on my hardware (p100 + 2xV100) and there are some pretty big differences between model families. Below are some of the results I got. Full details, some other models, including impacts to VRAM and scripts to run the benchmarks are on github here: https://github.com/bradrlaw/ai-server/blob/main/docs/benchmarking.md Gemma-31b scaled nicely with more n-max Qwen benefited most from a middle setting Everyone's favorite scaled well The smaller Gemma model behaved opposite of the larger one submitted by /u/bradrlaw [link] [comments]
Related
- Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti
- r/DestroyMyGame destroyed me to the void for using AI. I used Qwen 3.6 27B Q8 with MTP for about 20% of this single HTML file physics shooter game. I remember last year being blown away by GLM 4.5 Air being able to write a somewhat coherent HTML webpage.
Source: r/LocalLLaMA | 2026-07-24