Model Releases
the real unlock is not cheaper evaluation it is making specialized evaluators cheap enough to run continuously on production traces so evalu…
the real unlock is not cheaper evaluation it is making specialized evaluators cheap enough to run continuously on production traces so evaluation stops being a launch checklist and becomes a permanent
the real unlock is not cheaper evaluation it is making specialized evaluators cheap enough to run continuously on production traces so evaluation stops being a launch checklist and becomes a permanent feedback loop that actually improves agents over time 🚀Today we launched LangSmith Tuned Evaluators, starting with Perceived Error. Tuned Evaluators run on production traces to catch undesirable agent behavior and attach feedback that you can use in your agent improvement processes. In our benchmark, our tuned model beat frontier m…
Related
- 🚀Today we launched LangSmith Tuned Evaluators, starting with Perceived Error. Tuned Evaluators run on production traces to catch undesirabl…
- LangSmith’s new Tuned Evaluators look pretty interesting. Perceived Error can flag agent mistakes in production by picking up on user correc…
- New Tuned Evaluators from @LangChain Versioned judges you just 'turn on' in a tracing project The 'Perceived Error' Judge flags conversation…
- We built SmithDB: the database purpose built for agent observability workloads that now powers many parts of LangSmith. Agent observability …
Source: Harrison Chase (X) | 2026-08-18