Local Ai

Token/s down 50% after upgrade

Upgraded ollama from .32.6 to .32.15 ran a test on my qwen3:30b and it went from 90+ t/s down to 45+t/s System has 4 and v620 pro and one rtx3090 Before it was only using the amd, now it's using all 5

DGX agentreddit
local-air-ollama

Upgraded ollama from .32.6 to .32.15 ran a test on my qwen3:30b and it went from 90+ t/s down to 45+t/s System has 4 and v620 pro and one rtx3090 Before it was only using the amd, now it's using all 5 gpus and the token rate is awful, i rolled it back to .6 but it continues now to use all 5 gpus and only gets 45t/s... Is there an easy way to specify which model uses which gpus? Thanks. submitted by /u/Minister74 [link] [comments]

Source: r/ollama | 2026-08-25

Loading related sources…