Research
Data Compressibility Quantifies LLM Memorization
arXiv:2507.06056v4 Announce Type: replace Abstract: Large Language Models (LLMs) are known to memorize portions of their training data, sometimes even reproduce content verbatim when prompted appropri
arXiv:2507.06056v4 Announce Type: replace Abstract: Large Language Models (LLMs) are known to memorize portions of their training data, sometimes even reproduce content verbatim when prompted appropriately. Despite substantial interest, existing LLM memorization research has offered limited insight into how training data influences memorization and largely lacks quantitative characterization. In this work, we build upon the line of research that seeks to quantify memorization through data compressibility. We analyze why prior attempts fail to yield a reliable quantitative measure and show that a surprisingly simple shift from instance-level to set-level metrics uncovers a robust phenomenon, which we term the extit{Entropy--Memorization (EM) Linearity}. This law states that a set-level data entropy estimator exhibits a linear correlation with memorization scores.
Related
- Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
- Data Mixing for Large Language Models Pretraining: A Survey and Outlook
- Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
- Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
Source: arXiv cs.CL | 2026-04-21