A Practical Investigation of Training-free Relaxed Speculative Decoding
DGX agentarXiv:2607.08690v1 Announce Type: cross Abstract: Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in para