Model Releases

Single system with dual cards or two systems with single cards?

So I am in a conundrum and I'm thinking of asking for your opinion for the following: Currently, I have a 5800X3D gaming rig with a 7900XTX with its 24GB VRAM. It seems that for this subreddit, this c

DGX agentreddit
model-releasesr-localllama

So I am in a conundrum and I'm thinking of asking for your opinion for the following: Currently, I have a 5800X3D gaming rig with a 7900XTX with its 24GB VRAM. It seems that for this subreddit, this configuration seems to be GPU poor, judging from other's setups in here. :) I am actually eyeing to maybe get a AMD Radeon PRO v620 32GB, that would be used purely only for inference, as the 7900XTX is my main display card, so it's VRAM is always being used by the OS. The current card is a Sapphire 7900XTX Nitro+ Vapor-X and it's humongous. It is so large that its blocking the other PCIe slot, so I cannot actually slot another card as a second card in the motherboard. But I also have a smaller mini ITX system, that I use as my Docker server for my small homelab with Ubuntu 24.04. So here's my conundrum. Should I just slot the v620 into this second system and use it as a separate card, or should I get an open frame case for my main system, so that I can connect both cards with risers, so that I could get more combined VRAM across the cards? The former is much easier than the latter, of course, because I must essentially get a new frame case and gut my existing case and get a better PSU. Is it actually worth it to have a combined two-card system with 24+32GB VRAM, or just use them as separate systems? In your experience, have you used mixed cards and do they actually work combined like this? Currently the local "small SOTA" I run with my card, are Qwen3.6-27B & 35B and Gemma-4, all with Q4 quants. Having more VRAM in one system would would enable me to use better quantizations like Q6 or Q8, but would it using splitted across two cards on the PCIe bus. Would this make it slower, than the current 40-60 tps / 500pp I have with 27B on the single card? But if I would have a separate systems for these cards, I could maybe run Q5 quant on the v620 alone. Would this be good enough? Sorry for the thousand questions I ask. submitted by /u/noctrex [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-02

Loading related sources…