Agents
Learn more about the latest from @james_y_zou and our Frontier Agents Research team!
Learn more about the latest from @james_y_zou and our Frontier Agents Research team! To evaluate frontier AI agents, we need more complex tasks. But such tasks are also more prone to have design mista
Learn more about the latest from @james_y_zou and our Frontier Agents Research team! To evaluate frontier AI agents, we need more complex tasks. But such tasks are also more prone to have design mistakes. We audited 168 LLM/agent benchmarks—eg SWE-bench, Terminal-Bench 2—and found substantial issues in many tasks: ambiguous prompts, misaligned tests + other flaws…
Source: Together AI (X) | 2026-05-30