A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
DGX agentarXiv:2607.27735v1 Announce Type: new Abstract: Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafti