Tutorials
GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding
arXiv:2608.25343v1 Announce Type: new Abstract: Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated cor
arXiv:2608.25343v1 Announce Type: new Abstract: Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated correction pairs that are costly to maintain as query vocabularies evolve. Unsupervised correction with language models is attractive, yet in the short-query setting, unconstrained generation often over-corrects ambiguous inputs toward high-frequency phrases, causing intent drift. We propose extsc{GUIDE}, a generative unsupervised framework for CQC based on a confuse-then-clarify paradigm. extsc{GUIDE} encodes phonetically or visually confusable characters with shared-IDs and reconstructs the original query with an encoder--decoder architecture, which constrains correction to plausible confusion neighborhoods while learning from unlabeled query streams. A time-decayed, query-frequency-weighted objective further supports adaptation to rapidly changing query vocabularies. Experiments on extit{QSpell 250K} and a large-scale real-world dataset (extit{KwaiSearch}) show that extsc{GUIDE} consistently outperforms strong baselines, while online A/B testing further confirms gains in correction quality and downstream engagement.
Related
- Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing
- RCEM: Embedder Equipped with Query Rewriting Skill for Robust Conversational Search in Distributional Shift
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
Source: arXiv cs.CL | 2026-08-27