Safety
On The Statistical Limits of Self-Improving Agents
arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework
arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: under standard i.i.d. assumptions, distribution-free PAC learnability is preserved if and only if the policy-reachable family remains uniformly capacity-bounded. If reachable capacity can grow without bound, utility-rational self-changes can make learnable tasks unlearnable. We further introduce a simple Two-Gate guardrail -- a validation-improvement requirement plus a capacity cap -- that preserves this boundary and yields standard VC-rate guarantees. The broader implication is that self-modification must be constrained not only by objectives, but also by structural conditions that preserve the statistical prerequisites for learning. As AI systems become increasingly intelligent and autonomous, this framework provides a foundation for the statistical theory of self-improvement.
Related
- Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
- SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
- Self-Evolving Agents with Anytime-Valid Certificates
- MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
Source: arXiv cs.AI | 2026-08-12