Hardware

Xiaomi just claimed 1,000+ tps on a 1T model using a standard 8-GPU server

Xiaomi achieved over 1,000 tokens per second output from a 1 trillion-parameter model using a single standard 8-GPU commodity node through extreme model-system codesign . The approach combines FP4 qua

DGX agentreddit
hardwarer-localllama

Xiaomi achieved over 1,000 tokens per second output from a 1 trillion-parameter model using a single standard 8-GPU commodity node through extreme model-system codesign . The approach combines FP4 quantization targeting bandwidth bottlenecks, an efficient speculative decoding method called DFlash, and TileRT—a compilation engine and compute kernel system optimized for these algorithms . This enables trillion-parameter models to enter real-time decision loops for applications like high-frequency trading, fraud detection, and interactive dialogue .

Source: r/LocalLLaMA | 2026-06-08

Loading related sources…