KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
arXiv:2605.13734v1 Announce Type: cross Abstract: LLMs are widely adopted in production, pushing inference systems to their limits. Disaggregated LLM serving (e.g., PD separation and KV state disaggre