Model Releases

Running Qwen3.5-397B-A17B (4bit quants, 177 GB) on two DGX Sparks using llama.cpp with RPC and RDMA:

This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit

DGX agentx-post
model-releasesgeorgi-gerganov--x

This post documents a technical demonstration of running the large Qwen3.5-397B-A17B model across distributed hardware using llama.cpp with advanced networking protocols. The approach leverages 4-bit quantization to reduce the 177 GB model to a manageable size while utilizing RPC (Remote Procedure Call) and RDMA (Remote Direct Memory Access) for efficient inter-node communication between two DGX Spark systems.

Source: Georgi Gerganov (X) | 2026-04-27

Loading related sources…