QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
DGX agentarXiv:2604.08585v1 Announce Type: cross Abstract: Cache fusion accelerates generation process of LLMs equipped with RAG through KV caching and selective token recomputation, thereby reducing computati