Safety
> shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct…
> shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct > wrong tool called 3 steps earlier > no trace to follow >
shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct > wrong tool called 3 steps earlier > no trace to follow > no dataset to reproduce it > no eval to catch it next time > this wasn't a model problem > it was an ops problem > build, test, deploy, monitor - in that order > skip test: you deploy guesses > skip monitor: you lose the signal > skip the loop: every version starts from zero the fix: > small eval dataset before production > traces capturing the full agent trajectory > LLM-as-judge scoring tool calls and policy > feedback attached to the exact run that broke > shared infra so every team doesn't rebuild this Media
Source: Harrison Chase (X) | 2026-05-10