An Interpretable Latency Model for Speculative Decoding in LLM Serving
DGX agentarXiv:2605.15051v1 Announce Type: new Abstract: Speculative decoding (SD) accelerates large language model (LLM) inference by using a smaller draft model to propose multiple tokens that are verified b