Model Releases
How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent…
How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent's self-generated world knowledge actually improves its task
How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent's self-generated world knowledge actually improves its task success rate. The external guidance is then removed at inference. Result: A 14B model trained this way surpasses Gemini-2.5-Flash and gains +20% on WebVoyager and WebWalker. If agents can reliably improve themselves by exploring the world rather than waiting for human-labeled rewards, the bottleneck for scaling agentic systems shifts from data curation to environment design. Paper: https://arxiv.org/abs/2604.18131 Learn to build effective AI agents in our academy: https://academy.dair.ai/
Source: DAIR.AI (X) | 2026-04-22