Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
DGX agentarXiv:2605.14005v1 Announce Type: new Abstract: Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens a