KVBuffer: IO-aware Serving for Linear Attention
DGX agentarXiv:2605.19049v1 Announce Type: cross Abstract: Linear attention has recently gained significant attention for long-context inference due to its constant decoding cost with respect to context length