Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching
DGX agentarXiv:2606.07684v1 Announce Type: cross Abstract: Disaggregated serving alleviates memory bottlenecks in Large Language Model (LLM) inference but creates a severe communication bottleneck: transmittin