Model Releases
Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA:
This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit
This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit quantization to reduce the 177 GB model to a manageable size while utilizing RPC (Remote Procedure Call) and RDMA (Remote Direct Memory Access) for efficient inter-node communication between two DGX Spark systems.
Source: Georgi Gerganov (X) | 2026-04-27