Cartridges at Scale: Training Modular KV Caches over Large Document Collections
DGX agentarXiv:2606.04557v1 Announce Type: new Abstract: Large Language Models can reason over long contexts, yet prefilling millions of tokens is wasteful as much of the content remains static across queries.