SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering
DGX agentarXiv:2608.00311v1 Announce Type: new Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key