TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
DGX agentarXiv:2605.17170v1 Announce Type: new Abstract: Agentic workloads have emerged as a major workload for LLM inference. They differ significantly from chat-only workloads, requiring long-context process