Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA:
This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit