Research
Paper from Kimi: Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter [R]
Mooncake is the serving platform for Kimi developed by Moonshot AI, featuring a KVCache-centric disaggregated architecture that separates prefill and decoding clusters while leveraging underutilized C
Mooncake is the serving platform for Kimi developed by Moonshot AI, featuring a KVCache-centric disaggregated architecture that separates prefill and decoding clusters while leveraging underutilized CPU, DRAM, and SSD resources to implement a disaggregated KVCache pool. The system uses a KVCache-centric scheduler that balances throughput maximization with latency-related Service Level Objectives (SLOs). Mooncake achieves up to 525% throughput increase in simulated scenarios and enables Kimi to handle 75% more requests in real workloads.
Related
- ResBM: a new transformer-based architecture for low-bandwidth pipeline-parallel training, achieving 128× activation compression [R]
- Towards Faster Language Model Inference Using Mixture-of-Experts Flow Matching
Source: r/MachineLearning | 2026-04-18