Model Releases

1x32GB V100 vs 2x16GB V100 vs 5060ti 16GB for QWEN 3.8

Hi All, I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running llama.cpp on linux. Objective is to increase context

DGX agentreddit
model-releasesr-localllama

Hi All, I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running llama.cpp on linux. Objective is to increase context and use a better quant and also free up 5060 for other tasks. The options I am considering are 1x32GB V100 and 2x16GB V100. Theoretically, 2x16GB should be superior in terms of performance to 5060 and 1x32gb due to higher memory bandwidth. One issue I have to deal with is that I am limited in terms of CPU to GPU comms - I only have 2x x4 lines available. Any other good options in the same price range? UPDATE: Found a very interesting page, showing performance of multiple V100 with qwen 3.8 : https://domoticx.net/docs/llm-with-lama.cpp submitted by /u/ColorsOfCosmos [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-30

Loading related sources…