Hardware
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
arXiv:2604.18170v1 Announce Type: new Abstract: LLMs edit text and code by autoregressively regenerating the full output, even when most tokens appear verbatim in the input. We study Copy-as-Decode, a
arXiv:2604.18170v1 Announce Type: new Abstract: LLMs edit text and code by autoregressively regenerating the full output, even when most tokens appear verbatim in the input. We study Copy-as-Decode, a decoding-layer mechanism that recasts edit generation as structured decoding over a two-primitive grammar: references an input line range, ... emits new content. A token-level FSM guarantees syntactic validity, and a serving-layer primitive updates the KV cache for each copy span via a single parallel-prefill forward rather than N autoregressive steps -- sharing the parallel-forward kernel of speculative decoding but with input tokens as the draft and program-enforced acceptance replacing probabilistic verification. We report an upper-bound analysis that requires no end-to-end training. (i) Kernel speedup: on Qwen2.5-{1.5B, 7B}, copying N tokens via parallel prefill is 6.8imes--303imes faster than autoregressive (N in [8, 512], A100 80GB bf16). (ii) Copy ceiling: on ProbeEdit and HumanEvalPack-Fix (Py/JS), 74--98% of gold tokens are reachable under the line-level primitive; composed with the empirical kernel over each corpus's span histogram this yields a closed-form wall-clock bound of 29.0imes / 3.4imes / 4.2imes (13.0imes pooled). A token-level extension reaches 91--99% coverage with 4.5imes--6.5imes floors. (iii) Pipeline losslessness: oracle programs round-trip through the deterministic resolver on all 482 cases, localizing any downstream failure to span selection rather than the mechanism. A perturbation study shows pooled EM drops from 100% to 15.48% under off-by-one noise. A fine-tuning pilot on Qwen2.5-Coder-1.5B lifts HEvalFix-Py EM from 0/33 (untrained) to 12--17%, a learnability signal, not a production selector. Batched-serving integration and multi-file coverage are scoped as follow-up.
Related
- CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
- ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
- AdaSplash-2: Faster Differentiable Sparse Attention
- StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
Source: arXiv cs.CL | 2026-04-21