Agents
WANDR tests how well agents discover large sets of entities and verify specific facts about each one. It provides a dense, interpretable eva…
WANDR tests how well agents discover large sets of entities and verify specific facts about each one. It provides a dense, interpretable eval signal that reveals whether an agent fails, and where. The
WANDR tests how well agents discover large sets of entities and verify specific facts about each one. It provides a dense, interpretable eval signal that reveals whether an agent fails, and where. The pipeline also doubles as a semi-automated factory for training data.
Related
- AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
- How to evaluate AI agents, avoid reward hacking, and build better specs
- LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
Source: Perplexity (X) | 2026-07-14