Model Releases

AMD setup is fast with 200k ctx with Qwen 3.8, i didn't understand how/why?

Hi all, it is my very first post, because I am so confused. First my setup: I have 2 X 7900XTX with rocm7.2.1 and I use it to fine tune small models and local llm etc. I run Qwen 3.8 27b 8 bit version

DGX agentreddit
model-releasesr-ollama

Hi all, it is my very first post, because I am so confused. First my setup: I have 2 X 7900XTX with rocm7.2.1 and I use it to fine tune small models and local llm etc. I run Qwen 3.8 27b 8 bit version, When I limit ctx with proper number, Iike 184320, I get this result, super fast token reading but generation is 25 token/s. but you can see it is way different for not 200k ctx window. https://preview.redd.it/th2yns1zyikh1.png?width=537&format=png&auto=webp&s=822fadb49a3ee132ec860c5fdd5d7f773f943bff this is the result of 200k ctx window. I have 3x token generation, I didn't understand why 200000k is 3x faster, it is not even in binary system... https://preview.redd.it/qvza071txikh1.png?width=521&format=png&auto=webp&s=a64a5d9ed26476acddf997f5ff1f974906d696b0 I am so happy with 77 token/s generation, my agents are way faster witk 200k ctx model with it but I still don't know why... submitted by /u/Weary-Roof-1649 [link] [comments]

Related

Source: r/ollama | 2026-08-20

Loading related sources…