Still: Amortized KV Cache Compaction in a Single Forward Pass
DGX agentarXiv:2606.07878v1 Announce Type: new Abstract: The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call