Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding
DGX agentarXiv:2606.24957v1 Announce Type: new Abstract: While speculative decoding improves inference throughput for multi-batch long-context Large Language Models (LLMs), its efficiency is often limited by a