Applications
Excited to share our work on production-ready W4A8 inference, now integrated in vLLM! By combining 4-bit weights (low memory) with 8-bit act…
Excited to share our work on production-ready W4A8 inference, now integrated in vLLM! By combining 4-bit weights (low memory) with 8-bit activations (high compute), we hit the sweet spot for both deco
Excited to share our work on production-ready W4A8 inference, now integrated in vLLM! By combining 4-bit weights (low memory) with 8-bit activations (high compute), we hit the sweet spot for both decoding and prefill — up to 58% faster TTFT and 45% faster TPOT vs W4A16 on Hopper.
Related
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
- SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
Source: Cohere (X) | 2026-04-22