Research
ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization
arXiv:2608.12756v1 Announce Type: new Abstract: Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We int
arXiv:2608.12756v1 Announce Type: new Abstract: Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We introduce ReconSpan, which divides text into chunks that a backward decoder can reconstruct from a single contextual prefix code and retains one such code as the latent token for each chunk. The reconstruction criterion is applied when chunks are formed, allowing one trained autoencoder to produce average chunk lengths from 6.5 to 12.2. At matched average length, reconstruction-guided boundaries preserve more text than random boundaries. Readers of the resulting latent sequence recover topic information reliably but struggle to extract exact details.
Related
- CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding
- Incremental BPE Tokenization
- Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
Source: arXiv cs.CL | 2026-08-14