Model Releases
Need support for llama.cpp with multi GPU
Using llama.cpp I seem to be unable to get my to GPUs working tougether correclty, so I need help somehow. Setup: 96GB RAM, one Blackwell 5000 (48GB) and one 3090 (24GB). I am trying to run the UD-Q3_
Using llama.cpp I seem to be unable to get my to GPUs working tougether correclty, so I need help somehow. Setup: 96GB RAM, one Blackwell 5000 (48GB) and one 3090 (24GB). I am trying to run the UD-Q3_K_XL quant of Deepseek4 flash which has about 120GB size. Using just the Blackwell I would put most of the experts on the system RAM. I would have thought that putting some experts on the 3090 I would have thought to be faster (or at least not slower) than putting the experts just on the RAM. What I have tried: for all trails I try run with these common flags: --temp 1.0 --top-p 0.95 --min-p 0 -ngl 99 -ctk bf16 -ctv bf16 --jinja --flash-attn on --no-mmap baseline, using just the Blackwell (-dev CUDA0 --fit on --n-cpu-moe 32): full context and a generation speed of 16t/s that above but swapping "-dev CUDA0" with "--split-mode layer": fit fails trying to allocate 31gig on the 3090 adding -ts 2,1 results in exactly the same as above back to the drawing board, start with minimal setup using just the args from my first bullet point above): fit settles at 4096 context size trying to fix the context by setting -c 1048576 or fitc 1048576: fit now fucks up and just pushes 25 gigs on the blackwell and just 6 gigs on the 3090 Now I moved to manually setting the layers: agian the baseline with just the balckwell (-dev CUDA0 -ot "(1[2-9]|[2-9][0-9]).ffn_._exps.=CPU"): full context, same 16t/s speed as with moe-cpu just a above but without -dev CUDA0: speed drops to 11t/s manually puting expers on 3090 (-ot "(1[2-7]).ffn_._exps.=CUDA1" -ot "(1[8-9]|[2-9][0-9]).ffn_._exps.=CPU"): speed of 12t/s What am I doing wrong? submitted by /u/erazortt [link] [comments]
Related
- How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)
- NCCL-Free Tensor Parallelism on Dual Blackwell PCIe llama.cpp b9095 released!
- llama.cpp - NVFP4 native support on Blackwell from now - b8967
Source: r/LocalLLaMA | 2026-08-22