Research
Geometric Iterative Retrieval for Neural Audio Codec Resynthesis
arXiv:2608.19141v1 Announce Type: cross Abstract: Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generat
arXiv:2608.19141v1 Announce Type: cross Abstract: Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. Prior work has framed resynthesis as a choice between discrete token prediction and continuous regression. We argue that this dichotomy is incomplete and introduce geometric iterative retrieval, a paradigm that uses the RVQ layer hierarchy itself as a natural iterative decomposition in continuous codebook space. Rather than classifying over discrete vocabularies or regressing to a single target vector, our method performs contrastive retrieval in the codebook's geometric space. We evaluate our method on codec restoration tasks across speech and music, and show improvements over both single-pass token prediction and one-step regression baselines.
Related
- Exploring Token-Space Manipulation in Latent Audio Tokenizers
- LILAC: An Idempotent Neural Speech Codec
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
- Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
Source: arXiv cs.LG | 2026-08-20