Hardware
Unweight: how we compressed an LLM 22% without sacrificing quality
Running LLMs across Cloudflare’s network requires us to be smarter and more efficient about GPU memory bandwidth. That’s why we developed Unweight, a lossless inference-time compression system that ac
Running LLMs across Cloudflare’s network requires us to be smarter and more efficient about GPU memory bandwidth. That’s why we developed Unweight, a lossless inference-time compression system that achieves up to a 22% model footprint reduction, so that we can deliver faster and cheaper inference than ever before.
Related
- IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
- CASK: Core-Aware Selective KV Compression for Reasoning Traces
- CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
- NVIDIA NVbandwidth: Your Essential Tool for Measuring GPU Interconnect and Memory Performance
Source: Cloudflare AI | 2026-04-17