Model Releases
teams want to understand what their agents are doing but it’s hard to sift through all that trace data at scale but you can basically find a…
teams want to understand what their agents are doing but it’s hard to sift through all that trace data at scale but you can basically find any signal by fine-tuning a small, cheap model on data showin
teams want to understand what their agents are doing but it’s hard to sift through all that trace data at scale but you can basically find any signal by fine-tuning a small, cheap model on data showing what that signal looks like we launched in Beta a couple months ago and got great feedback on teams being able to cheaply label ever single trace with a judge they care about the future is hundreds of small evaluations running constantly, ultra cheaply, processing every trace data mining applied at massive scale to understand then improve agents over time if that future is interesting to you’d reach out, we’d love to help you build these! Introducing LangSmith Tuned Evaluators They automatically score agent behavior in production, starting with Perceived Error. Perceived Error is one of the clearest signals that your agent is giving users a helpful experience. In our benchmark, our specialized model outperformed e…
Related
- Got this setup in Claude Code to trace models routed through @merge_api! Being able to trace what your agents are doing is so valuable, I lo…
- I detected a bad Agent action, what do I do about it? this is pretty much the main question that will power the future’s Human+Agent driven …
- If you are building real-time voice agents with @GoogleDeepMind Gemini Live, you can now trace your speech-to-speech agent loops directly in…
Source: Harrison Chase (X) | 2026-08-18