Model Releases
brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and …
brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and encounter. They’re not really random benchmark tasks, they r
brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and encounter. They’re not really random benchmark tasks, they reflect our priors on likely user behaviors to make sure the agent passes those cases in the product before we just ship it 2. With that said, the best Evals often aren’t made from scratch, they’re discovered from real world Traces. I’ll literally never ship a perfect agent first try, we need user feedback and failures from Traces to make Evals and then make sure these errors don’t happen again 3. At a basic level, Evals give us some measurable, apples to apples way of comparing performance. Ex: is my agent good today and also in 1 month when I go to try a new model? 4. Evals ~= Environments, we need some place to run Eval Tasks, which is defined by the environment setup. This should mirror prod as much as possible. The more that Eval drifts from prod, the higher my Sim2Real gap and the less I can trust the numbers 5. Evals are our regression tests. Sometimes a prompt change might fix something today while breaking something I changed last week. Evals help us catch that 6. Evals are our training data. They map out what we hope the agent should be & do. We literally fit the agent to evals in hopes of making more Evals pass. Good evals —> good agent. In some rough way: Agent = fit(model, evals) 7. It’s ok if some Evals fail today if it means this Eval is simply too hard for today’s models, but it’s something to strive towards for the next gen of agents. But you should still do agent engineering today to make them pass if you can, it’s a goal to engineer towards 8. The best Eval is an Eval that actually exists. I think still today Evals are daunting because of the blank canvas problem. “Where do I even start??” But I find small Unit Test style evals are a great place to start to feel like I’m building momentum. I can get a real number + Unit Tests are familiar to many folks 9. Evals need to undergo spring clean. Not every is relevant over time. Models get smarter, agent priorities shift, user behavior shifts. Just how we clean up dead code, we clean up dead evals. This saves us money and avoids us training in agent behaviors we don’t care about anymore we make Evals not because it’s what everyone says to do, but because it’s one of the most meaningful levers we can pull in building better agents. and if we can help make that easier or at the very least make it less scary to start then that’s great :)
Source: Harrison Chase (X) | 2026-05-20