Agents
AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
arXiv:2608.26004v1 Announce Type: cross Abstract: Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control laten
arXiv:2608.26004v1 Announce Type: cross Abstract: Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an asymmetric speculative decoding framework that breaks this symmetry: a lightweight drafter reads the full input while the large verifier operates on the compressed view. The drafter steers the verifier via a contrastive elta-fusion of logits, modulated by a divergence-aware acceptance gate that preserves verification stability and high draft acceptance rates. Evaluated across four agentic capabilities and two end-to-end agent benchmarks, AsymSpec reaches approx 90% of full-context accuracy on average, delivering 1.3--1.7imes throughput speedups at 0.2--0.3imes the compute cost on isolated text capabilities. These results show that asymmetric context access yields substantial gains precisely when compression discards critical reasoning signals.
Related
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
- HaS: Accelerating RAG through Homology-Aware Speculative Retrieval
- ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
- Mitigating Context Interference for Reliable and Efficient Search Agents
- A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
Source: arXiv cs.CL | 2026-08-27