Hardware
We are excited to have @baseten as a day 0 launch partner for Kimi K2.6! Their inference stack brings KV-aware routing, NVFP4 on Blackwell, …
We are excited to have @baseten as a day 0 launch partner for Kimi K2.6! Their inference stack brings KV-aware routing, NVFP4 on Blackwell, multi-modal hierarchical caching, and prefill-decode disaggr
We are excited to have @baseten as a day 0 launch partner for Kimi K2.6! Their inference stack brings KV-aware routing, NVFP4 on Blackwell, multi-modal hierarchical caching, and prefill-decode disaggregation, so K2.6 runs the way it's meant to in production. Try it out at: https://www.baseten.co/library/kimi-k26 Kimi K2.6 has landed, and it is live on Baseten! We have baked in multiple inference optimizations so that you can leverage Kimi K2.6 in production right away. To run Kimi K2.6, Baseten uses: -> The Baseten Inference Stack with advanced optimizations, including KV-aware routing -…
Related
- StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
- CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
- Nvidia published DWDP (Distributed Weight-Data Parallelism), a new inference parallelism strategy focused on prefill. It sounds slightly ins…
- Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
Source: Kimi/Moonshot (X) | 2026-04-20