SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
DGX agentarXiv:2506.01979v4 Announce Type: replace-cross Abstract: Recently, speculative decoding (SD) has emerged as a promising technique to accelerate LLM inference by employing a small draft model to propo