Model Releases
Getting a second GPU in addition to my RTX3090
Hello, I've been learning how to use local LLMs for a year or so on my workstation, using a RTX3090. Current setup : - i5 12400 - 64gb RAM - RTX 3090 - OS : Fedora KDE workstation I'm using LMStudio t
Hello, I've been learning how to use local LLMs for a year or so on my workstation, using a RTX3090. Current setup : - i5 12400 - 64gb RAM - RTX 3090 - OS : Fedora KDE workstation I'm using LMStudio to serve mainly these models : - Qwen 3.5 27B - Gemma 4 26b a4b And then VSCode + kilocode / Continue for light coding / scripting / log analysis tasks (i'm a sysadmin). It works very well, but i'm hitting the context size ceiling quite fast with this setup. This prevents me to work on bigger projects. My goal is to buy another GPU to provide more VRAM, and since i'm using linux, i'd prefer an AMD GPU. A 16gb Radeon 9070 would be nice in this regard : it's natively supported on Linux, doesn't cost an arm, it's powerful enough for casual gaming, and doesn't have crazy power requirements. I've read here and there that mixing AMD and Nvidia is now well supported (https://www.reddit.com/r/LocalLLaMA/comments/1qea29t/mix_of_amd_nvidia_gpu_in_one_system_possible/), especially with LMStudio. Since this is moving and evolving very quickly, how are the support and the performance of mixing a 24gb nvidia gpu + 16gb AMD gpu in 2026 ? Is it a better idea to go for another Nvidia GPU instead ? From my understanding, the idea would be to load the model in the 24Gb 3090, and then reserve the 16gb GPU for context. Correct me if I'm wrong. I've read multiple times 36 / 40 Gb VRAM is the sweet spot for the models i'm using (at q4 quant). Also, adding another 3090 is a solution i'd like to avoid because : - it's expensive and hard to find in good condition at the moment - adding another 350W and 3 slot GPU would be a challenge to cool down in my PC case - I probably would have to buy a new and bigger PSU too. So, is adding a 16gb AMD GPU to my current setup a good idea ? Thanks ! submitted by /u/kobbalt [link] [comments]
Related
- NCCL-Free Tensor Parallelism on Dual Blackwell PCIe llama.cpp b9095 released!
- Made a program using LocalLLM based on llama.cpp for fellow Book Lovers!
- FYI You dont need expensive networking for multi-node gpu. 30t/s laguna Q2_K_XL (39.7GB) on 2x4060+1x4060 using a $20 usb->ethernet.
Source: r/LocalLLaMA | 2026-07-25