Tutorials

GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding

arXiv:2608.25343v1 Announce Type: new Abstract: Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated cor

DGX agentpaper
tutorialsarxiv-cs-cl

arXiv:2608.25343v1 Announce Type: new Abstract: Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated correction pairs that are costly to maintain as query vocabularies evolve. Unsupervised correction with language models is attractive, yet in the short-query setting, unconstrained generation often over-corrects ambiguous inputs toward high-frequency phrases, causing intent drift. We propose extsc{GUIDE}, a generative unsupervised framework for CQC based on a confuse-then-clarify paradigm. extsc{GUIDE} encodes phonetically or visually confusable characters with shared-IDs and reconstructs the original query with an encoder--decoder architecture, which constrains correction to plausible confusion neighborhoods while learning from unlabeled query streams. A time-decayed, query-frequency-weighted objective further supports adaptation to rapidly changing query vocabularies. Experiments on extit{QSpell 250K} and a large-scale real-world dataset (extit{KwaiSearch}) show that extsc{GUIDE} consistently outperforms strong baselines, while online A/B testing further confirms gains in correction quality and downstream engagement.

Related

Source: arXiv cs.CL | 2026-08-27

Loading related sources…