Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
DGX agentarXiv:2607.27269v1 Announce Type: new Abstract: Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (