Practical Online KV Cache Compaction for LLM Agents: An Empirical Study
DGX agentarXiv:2608.00902v1 Announce Type: new Abstract: LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV c