SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance
DGX agentarXiv:2606.09441v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) injects LLM queries with relevant documents to improve response quality. This injection increases prompt length and