Local Ai
Token/s down 50% after upgrade
Upgraded ollama from .32.6 to .32.15 ran a test on my qwen3:30b and it went from 90+ t/s down to 45+t/s System has 4 and v620 pro and one rtx3090 Before it was only using the amd, now it's using all 5
Upgraded ollama from .32.6 to .32.15 ran a test on my qwen3:30b and it went from 90+ t/s down to 45+t/s System has 4 and v620 pro and one rtx3090 Before it was only using the amd, now it's using all 5 gpus and the token rate is awful, i rolled it back to .6 but it continues now to use all 5 gpus and only gets 45t/s... Is there an easy way to specify which model uses which gpus? Thanks. submitted by /u/Minister74 [link] [comments]
Source: r/ollama | 2026-08-25