DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting
DGX agentarXiv:2607.07409v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel. Block-parallel drafters such as DFlash furthe