Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
DGX agentarXiv:2511.22972v3 Announce Type: replace Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but suffer from high inference latency due to their autoregressive gene