Agents
a very cool Harbor x LangSmith flow I love to help you “look at the data”: 1. you do evals or rollouts for RL 2. all reward metrics and trac…
a very cool Harbor x LangSmith flow I love to help you “look at the data”: 1. you do evals or rollouts for RL 2. all reward metrics and traces and rollouts get automatically propulates into Experiment
a very cool Harbor x LangSmith flow I love to help you “look at the data”: 1. you do evals or rollouts for RL 2. all reward metrics and traces and rollouts get automatically propulates into Experiments & Tracing projects 3. we send agents (or use Engine) to understand how training or harness engineering affects agent behavior over each iteration Evals ~= Environments! more traditional Evals will become environment shaped as teams measure their complex agents rigorously evaluating agents in a way that reflects prod means creating an environment that mirrors prod late last year, we went all in on Harbor for our open source evals! and now anyone can also do the same and get all of that neatly organized to view across your team in LangSmith for both our Eval runs and RL rollouts, it’s very cool to have a shared source of truth where the team can understand how each targeted experiment affects our agent
Source: Harrison Chase (X) | 2026-06-30