Agents
Read more on what we've learned while developing long-horizon agent evals:
Read more on what we've learned while developing long-horizon agent evals: Curious finding while creating evals and benchmarks for long-horizon (100+ turn) agents While it’s generally thought that a d
Read more on what we've learned while developing long-horizon agent evals: Curious finding while creating evals and benchmarks for long-horizon (100+ turn) agents While it’s generally thought that a direct swap to open source models can bring immediate cost savings, that’s not what we saw off the bat. Two factors play a major role 👇
Source: Harrison Chase (X) | 2026-05-19