Model Releases

teams want to understand what their agents are doing but it’s hard to sift through all that trace data at scale but you can basically find a…

teams want to understand what their agents are doing but it’s hard to sift through all that trace data at scale but you can basically find any signal by fine-tuning a small, cheap model on data showin

DGX agentx-post
model-releasesharrison-chase--x

teams want to understand what their agents are doing but it’s hard to sift through all that trace data at scale but you can basically find any signal by fine-tuning a small, cheap model on data showing what that signal looks like we launched in Beta a couple months ago and got great feedback on teams being able to cheaply label ever single trace with a judge they care about the future is hundreds of small evaluations running constantly, ultra cheaply, processing every trace data mining applied at massive scale to understand then improve agents over time if that future is interesting to you’d reach out, we’d love to help you build these! Introducing LangSmith Tuned Evaluators They automatically score agent behavior in production, starting with Perceived Error. Perceived Error is one of the clearest signals that your agent is giving users a helpful experience. In our benchmark, our specialized model outperformed e…

Related

Source: Harrison Chase (X) | 2026-08-18

Loading related sources…