SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
arXiv:2605.02888v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference by using a small draft model to propose candidate tokens that a larger target mo