Tutorials
KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
arXiv:2604.12397v1 Announce Type: new Abstract: Standard Large Language Model (LLM) pre-training typically treats corpora as flattened token sequences, often overlooking the real-world context that hu
arXiv:2604.12397v1 Announce Type: new Abstract: Standard Large Language Model (LLM) pre-training typically treats corpora as flattened token sequences, often overlooking the real-world context that humans naturally rely on to contextualize information. To bridge this gap, we introduce Knowledge Coordinate Conditioning (KoCo), a simple method that maps every document into a three-dimensional semantic coordinate. By prepending these coordinates as textual prefixes for pre-training, we aim to equip the model with explicit contextual awareness to learn the documents within the real-world knowledge structure. Experiment results demonstrate that KoCo significantly enhances performance across 10 downstream tasks and accelerates pre-training convergence by approximately 30%. Furthermore, our analysis indicates that explicitly modeling knowledge coordinates helps the model distinguish stable facts from noise, effectively mitigating hallucination in generated outputs.
Related
- Learning is Forgetting: LLM Training As Lossy Compression
- What do Language Models Learn and When? The Implicit Curriculum Hypothesis
- Linear Representations of Hierarchical Concepts in Language Models
- Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
Source: arXiv cs.CL | 2026-04-15